Ollama's Apple Silicon Push, Gemma 3 Fixes, and Cloud Llama
Tools
Ollama MLX Support for Local VS Code Copilot
Run ollama start inside VS Code Copilot Chat to run models locally with MLX. This is huge for Apple Silicon users: you get the GUI and workflow of VS Code with the speed and privacy of local inference. No API keys, no cloud latency.
Gemma 3 Checkpoint Fix in vLLM
Checkpoint loading for Gemma3 was breaking in the previous engine; this patch usually fixes the runtime errors immediately. If you use vLLM for vision-language tasks, updating now saves you a debug cycle.
News
LaTeX Verification for Qwen3 Math Models
Using a real LaTeX compiler to verify that the model's mathematical reasoning is actually correct. This cuts down on the hallucinations in your draft generation pipelines.
DeepSeek R1 Continues to Update on HuggingFace
DeepSeek is dropping the actively maintained weights while other providers are tightening their grip. This is your signal to build against a stable, high-quality baseline that won't get deprecated tomorrow.
Resources
Guide: Preventing LLMs from 'Leaking' Old Context
This is a curated list of models that effectively 'forget' context outside a specific window, plus the right configurations to use actually new tokens. Essential bookmark for anyone filling a context window with data.
Stay Ahead
Delivered each morning.