DeepSeek V4-Flash brings frontier agent work to bargain pricing
The Neuron ● Covered by 5 sources
DeepSeek just pushed V4-Flash into public beta with a big jump in agent skills. It's cheap, it's fast, and it's now genuinely good at running tools and code on its own.
DeepSeek's changelog reads like a diary of relentless iteration, and the July 31 entry stands out from the rest. V4-Flash-0731 is now live in public beta, and the benchmark jumps are the kind that make you re-read the numbers twice. Terminal Bench 2.1 hit 82.7, Cybergym landed at 76.7, and DSBench-FullStack came in at 68.7 — all comfortably ahead of the earlier V4-Pro-Preview. For a model marketed as the budget tier, that's a strange place to be beating your own flagship.
What makes this update notable isn't a new architecture. DeepSeek says V4-Flash-0731 keeps the exact same size and structure as the preview version; the gains come purely from re-post-training. That's a quiet flex. It suggests DeepSeek thinks it still has a lot of headroom to extract from existing weights before it needs to spend compute on a bigger model, which cuts against the industry's usual instinct to scale up every six months.
The agent-specific numbers matter more than the headline scores. Toolathlon verified sits at 70.3, and NL2Repo — a test of turning plain instructions into working repositories — reaches 54.2. These are the benchmarks that predict whether a model can actually run a coding agent unsupervised for an afternoon rather than just answer trivia well. DeepSeek has also built native support for the Responses API format and tuned the model specifically for Codex, which tells you exactly who it's chasing: developers who want an agentic coding assistant without OpenAI or Anthropic pricing.
The rest of the changelog is a reminder of how fast this company iterates. Legacy names like deepseek-chat and deepseek-reasoner are being retired in three months, folded into V4-Flash's thinking and non-thinking modes. V4-Pro, meanwhile, is untouched for now, with a full release
My take
I run a site that lives and breathes AI news, and even I find DeepSeek's release cadence exhausting to track — which is exactly the point. They're optimizing for developer mindshare through sheer frequency and cheap agentic capability, not headline-grabbing scale, and that strategy is quietly working better than most closed labs want to admit.
Read more about this at: The Neuron