LLM News Digest

Curated updates on local AI models and edge computing

Executive Summary

The local LLM landscape is heating up: vLLM v0.25.0 and Ollama v0.32.5 bring major performance boosts for quantized model inference. Tether QVAC released open-weight edge-optimized vision-language models for consumer smartphones, while MSI launched a dedicated PRO MAX EDGE AI+ desktop for offline 120B-parameter LLM deployment. On the hardware front, RTX 5090 benchmarks reveal that short generation, long context, and code tasks each need separate speed measurements for accurate sizing.

vLLM v0.25.0 & Ollama v0.32.5 Boost Local AI Performance

PatentLLM Blog

vLLM v0.25.0 brings improved foundational LLM inference with enhanced quantized-model support, while Ollama v0.32.5 adds performance optimizations for local model runners. Together they represent significant upgrades for the open-source local AI ecosystem.

Local LLM 2026: Every Major Model Release + Ollama Status

PromptQuorum

As of July 2026, ten releases have moved past the Q1 2026 wave and now define the current state of the art for locally-runnable models, led by GLM-5.2 for raw capability, Kimi K2.6 and Laguna XS 2.1 for coding, and gpt-oss:20b for small-footprint deployment.

The Best Open Source and Open-Weight LLM Models to Run Locally in 2026

Hugging Face Blog

Hugging Face's curated guide covers the top open-weight models suitable for local deployment as of July 2026. Alibaba's Qwen3 is highlighted as the leading default choice for most developers, with DeepSeek R1 and Kimi K2.6 also featured for strong reasoning and capability.