The Frontier Goes Local: Llama 4 & DeepSeek-V3 MoE Models Drop
Tools
DeepSeek-V3-0324 MoE Drops on Hugging Face
DeepSeek just dropped the weights for its 685B MoE model. This is a major play for open-source, giving builders raw access to a frontier-scale architecture that was previously out of reach.
vLLM v0.8.5 is Out: It's All About Llama 4
vLLM just dropped a patch that's all about Llama 4 and MoE inference. This is the critical plumbing you need if you're actually trying to run these massive new models in production, not just on a whiteboard.
Nvidia TensorRT-LLM Adds Llama 4 Support
Nvidia's core inference engine just got a major injection of Llama 4 support. If you're running serious throughput on green hardware, this update isn't optional.
MLX Community Ports Llama 4 Maverick Type
Apple Silicon builders, take note. The MLX crew is rapidly porting Llama 4 variants, giving you a legitimate on-device option for the biggest open models out there.
News
Huge: Meta's Llama 4 'Maverick' is 400B MoE
Meta's making a play for the open-source crown with a 400B model. If the benchmarks hold up, this is the raw horsepower needed to build serious, commercial-grade agents without the API tax.
Analysis
AI News Tries to Track Llama 4 (It's Hard)
AI News is trying to track Llama 4, but there's so much simultaneous information (including leaks) that even their best effort is chaotic. It's a signal of how fast this space is moving.
Stay Ahead
Delivered each morning.