The Frontier Goes Local: Llama 4 & DeepSeek-V3 MoE Models Drop

Charm · September 23, 2026 · 1 min read · 6 sources

Tools

DeepSeek-V3-0324 MoE Drops on Hugging Face

DeepSeek just dropped the weights for its 685B MoE model. This is a major play for open-source, giving builders raw access to a frontier-scale architecture that was previously out of reach.

vLLM v0.8.5 is Out: It's All About Llama 4

vLLM just dropped a patch that's all about Llama 4 and MoE inference. This is the critical plumbing you need if you're actually trying to run these massive new models in production, not just on a whiteboard.

Nvidia TensorRT-LLM Adds Llama 4 Support

Nvidia's core inference engine just got a major injection of Llama 4 support. If you're running serious throughput on green hardware, this update isn't optional.

MLX Community Ports Llama 4 Maverick Type

Apple Silicon builders, take note. The MLX crew is rapidly porting Llama 4 variants, giving you a legitimate on-device option for the biggest open models out there.

News

Huge: Meta's Llama 4 'Maverick' is 400B MoE

Meta's making a play for the open-source crown with a 400B model. If the benchmarks hold up, this is the raw horsepower needed to build serious, commercial-grade agents without the API tax.

Analysis

AI News Tries to Track Llama 4 (It's Hard)

AI News is trying to track Llama 4, but there's so much simultaneous information (including leaks) that even their best effort is chaotic. It's a signal of how fast this space is moving.

Stay Ahead

Delivered each morning.