Curated updates on local AI models and edge computing
The local LLM landscape is heating up: vLLM v0.25.0 and Ollama v0.32.5 bring major performance boosts for quantized model inference. Tether QVAC released open-weight edge-optimized vision-language models for consumer smartphones, while MSI launched a dedicated PRO MAX EDGE AI+ desktop for offline 120B-parameter LLM deployment. On the hardware front, RTX 5090 benchmarks reveal that short generation, long context, and code tasks each need separate speed measurements for accurate sizing.
vLLM v0.25.0 brings improved foundational LLM inference with enhanced quantized-model support, while Ollama v0.32.5 adds performance optimizations for local model runners. Together they represent significant upgrades for the open-source local AI ecosystem.
As of July 2026, ten releases have moved past the Q1 2026 wave and now define the current state of the art for locally-runnable models, led by GLM-5.2 for raw capability, Kimi K2.6 and Laguna XS 2.1 for coding, and gpt-oss:20b for small-footprint deployment.
Hugging Face's curated guide covers the top open-weight models suitable for local deployment as of July 2026. Alibaba's Qwen3 is highlighted as the leading default choice for most developers, with DeepSeek R1 and Kimi K2.6 also featured for strong reasoning and capability.