KVC swapp、Edge MoE とスパースカーネルの最前線
Tools
llama.cpp Launches Bounty Blitz for Mature Vulkan Support
The project is offering over $50k in bounties to fix the Vulkan backend across different hardware configurations. If you have a non-NVIDIA GPU sitting in a drawer, this is your chance to get paid to make it actually work.
vLLM 0.8.5 Introduces Runtime KV Cache Swapping
vLLM 0.8.5 sneaks in a feature that lets you cache and reuse KV caches on the fly. For builders running multiple agents sharing the same context, this is a massive leverage point for cost and latency reduction.
Apple’s MLX 0.25 Ships With Native Sparse Kernel Support
MLX Version 0.25 just dropped with native sparse kernel support, allowing for much cheaper inference on Apple Silicon. If you're optimizing for the M-series, this update lets you push the hardware limits far beyond the Metal defaults.
Analysis
Baidu Open-Sources ChunkFormer for Streaming ASR
This analysis breaks down the architectural shift from dense transformers to sparse linear-recurrence blocks. If you're looking for the next wave of efficient, runnable models after Mamba, this paper points to the exact direction the research is heading.
News
Qwen Drops Qwen3-30B-A3B, Making MoE Viable at the Edge
Qwen has released weights for a 1.5B parameter MoE model that claims to outperform Qwen2.5-7B. This is a perfect candidate for on-device deployment where VRAM is tight but high reasoning capability is non-negotiable.
Stay Ahead
Delivered each morning.