Carl Franzen writes that DeepSeek has launched V4.1-Flash, a model featuring a 552-billion-parameter mixture-of-experts backbone designed to drastically reduce costs for long-context workflows through specialized caching and architecture. The model offers extremely low off-peak rates of $0.003 per million cached input tokens, making it highly competitive against frontier models like GPT-5.6 Sol and Claude Opus 5 when used in repetitive agentic loops. While its total parameter count has increased significantly compared to previous versions, its Causal Encoder-Decoder architecture aims to minimize compute requirements during the prefill stage of inference.
- V4.1-Flash utilizes a Causal Encoder-Decoder architecture that activates 8 billion parameters during prefill and 16 billion during decoding.
- The model features an extremely high context window of up to 1 million tokens.
- DeepSeek's technical report highlights the use of FP4 KV caching, which reduces global KV cache size by approximately one-quarter compared to its predecessor.
- Off-peak hours for lower pricing are scheduled from Monday through Friday, specifically between 01:00โ04:00 UTC and 06:00โ10:00 UTC.