DeepSeek released the DeepSeek-V4.1-Flash multimodal Mixture-of-Experts model, including long-context serving optimizations and updated API routing and pricing
Model release Provisional 84% confidence first seen
DeepSeek announced and released the DeepSeek-V4.1-Flash model, highlighting long-context performance improvements such as a reduced KV-cache footprint and serving-oriented techniques. Separate coverage also described a change to API routing and billing in which DeepSeek-V4.1-Flash would serve requests and set pricing for V4-Pro temporarily, with older variants being retired.