TLDRocket
Sign in

DeepSeek releases V4.1-Flash, says it outperforms flagship V4-Pro

SiliconANGLE Duncan Riley Covered by 2 sources

DeepSeek just swapped its V4-Pro API traffic to a smaller Flash model. It says the new model is faster, cheaper, and even beats the flagship on some tests.

Based on reporting by SiliconANGLE, Duncan Riley — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

DeepSeek has released V4.1-Flash, the smallest model in a new family, and it’s making a pretty bold claim: the smaller one beats the bigger one. The Chinese startup said tests by multiple parties put the open-weight model ahead of DeepSeek-V4-Pro on performance, cost, speed, and total runtime.

The company is backing that claim with a live product move, not just a blog post. Starting Sept. 14, API requests that used to go to V4-Pro will be answered by V4.1-Flash instead, and billed at the smaller model’s rates until a V4.1-Pro version arrives. DeepSeek also retired V4-Flash and the experimental vision model it shipped in August, so calls to either now land on V4.1-Flash.

Under the hood, this is a mixture-of-experts model with 552 billion parameters, up from 284 billion in V4-Flash. But DeepSeek says only 8 billion parameters stay active while it processes a prompt, and 16 billion while generating output. Image understanding, which was only in that experimental release last month, is now part of the model itself.

A lot of the work went into shrinking the key-value cache. DeepSeek says it stores those entries in four-bit floating-point format, bringing the global footprint to 890 bytes per token, about a quarter of what V4-Flash needed. Persistent cache storage on SSDs is also said to fall to about an eighth of the previous generation.

On its own benchmark table, DeepSeek compares V4.1-Flash at maximum reasoning effort with Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.6 Sol. It scored 90.6 on Terminal-Bench 2.1, just ahead of Opus 5 at 89.1 and GPT-5.6 Sol at 88.8. On DeepSWE v1.1, it solved 74.2% of tasks, compared with 74% for Opus 5 and 62.7% for V4-Pro. Both U.S. models still lead on GPQA Diamond, though.

The pricing pitch is hard to miss. Off-peak API use is 15 cents per million uncached input tokens and 60 cents per million output tokens, with rates doubling during weekday peak windows. DeepSeek says developers still using V4-Pro would pay $3.96 per million output tokens at peak, versus $1.20 for V4.1-Flash, which works out to roughly a 70% cut. The weights are on Hugging Face under the MIT license, and the model is already live in DeepSeek’s web and mobile apps.

My take — AI-written commentary, not fact-checked reporting

DeepSeek is doing what the industry loves to pretend it invented: making smaller look smarter. If a cheaper model can take over traffic from the flagship and still brag about benchmark wins, that’s not a side quest — that’s the whole business model. The awkward part is that this all lands the same day Anthropic names DeepSeek in a distillation complaint, which is a very on-brand way for AI to say hello.

Read more about this at: SiliconANGLE

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.