TLDRocket
Sign in

deepseek-ai/DeepSeek-V4-Flash-0731

Simon Willison's Weblog Simon Willison Covered by 6 sources

DeepSeek dropped a new 304-billion-parameter model called V4 Flash, and it's absurdly cheap to run. It's beating models nearly 50% bigger while costing pennies per million tokens.

Based on reporting by Simon Willison's Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

DeepSeek just quietly pushed out V4 Flash-0731, the newest member of its V4 family, and the headline isn't the size — it's the price-to-performance ratio. At 304 billion parameters, weighing in around 167GB on Hugging Face, it's not a small model by any stretch. But Artificial Analysis has it ranked ahead of MiniMax M3, a competitor sitting at 428 billion parameters. That's a lighter model beating a heavier one on their intelligence benchmarks, which is the kind of result that makes you double-check the numbers.

Then there's the pricing: $0.14 per million input tokens and $0.27 per million output tokens. Put that next to what the big labs charge for comparable capability and it starts to look like DeepSeek is giving the intelligence away. On the Intelligence Index versus Cost per Intelligence chart, V4 Flash sits in a spot that most vendors would love to occupy — high capability, rock-bottom cost.

Simon Willison ran his usual pelican-riding-a-bicycle test on it through OpenRouter and got mixed results depending on settings. With the default reasoning level, the output was mediocre — a forgettable pelican. But cranking the reasoning effort up to "high" via the CLI (llm -m openrouter/deepseek/deepseek-v4-flash-0731 -t pelican -o reasoning_effort high) produced a noticeably better result. It's a small but telling detail: this model's agentic and reasoning improvements seem to actually require you to ask for them, rather than showing up by default.

That tuning knob matters more than it sounds. A lot of "flash" or fast-tier models trade reasoning depth for speed and cost, and users who never touch the settings end up judging the model on its weakest configuration. DeepSeek's positioning here — cheap by default, genuinely capable when you dial it up — is a smart way to court both casual API users and people building serious agentic pipelines who'll happily pay the reasoning-effort tax in latency for better output.

The broader story is DeepSeek's continued habit of releasing frontier-adjacent models at prices that make Western labs' pricing pages look almost embarrassing. Whether that's subsidized, strategic, or just reflects genuinely lower training and inference costs in China, the effect on the market is the same: it resets what

My take — AI-written commentary, not fact-checked reporting

I'll say what I keep saying: DeepSeek's real product isn't the model, it's the price pressure it puts on everyone else's roadmap. Open-weight releases at these price points are the best thing happening in AI right now, and every closed-lab exec quietly hates that a 304B model with a $0.14 input price is embarrassing their margin slides. The reasoning-effort dependency is the one asterisk worth remembering before you crown anything the value champion.

Read more about this at: Simon Willison's Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.