Bron
Het verhaal
On 10 Sep 2026 DeepSeek officially released DeepSeek-V4.1-Flash, billed as the smallest model in its new architecture family with native visual understanding. It is a 552B-parameter multimodal MoE (Causal Encoder–Decoder): ~8B active parameters on input prefill and ~16B on decode, with up to a 1M-token context and a much smaller KV cache (~1/4 HBM and ~1/8 SSD vs V4-Flash per DeepSeek). Weights are on Hugging Face under MIT; the API model name is deepseek-flash. Legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash (prior Flash/Vision-Exp retired). DeepSeek says third-party tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total time, and from 12:00 Beijing on 14 Sep 2026 until a future V4.1-Pro, deepseek-v4-pro requests will route to V4.1-Flash at Flash pricing.