Local LLM Hub

News · Games · Tools — all running locally


Sunday, August 23, 2026 · Updated at 12:44 PM PDT

Local LLM tooling keeps maturing fast: Ollama 0.30 shipped with improved GGUF/llama.cpp compatibility alongside its MLX engine, while both Ollama and LM Studio added Anthropic-compatible endpoints letting local models drop into agent workflows. On-device AI is spreading too — AI browsers like Puma run Qwen/Gemma fully offline on phones, and NPU advances (Qualcomm/CXMT 3D DRAM, VeriSilicon 40+ TOPS IP) push billion-parameter models to real-time speeds.

Ollama 0.30 Anthropic API On-Device AI NPU Advances