DeepSeek’s smaller model just outperformed its own flagship
The New Stack 3 weeks ago 26 ● 6 sources
DeepSeek released V4-Flash-0731, a smaller model with 284 billion total parameters that outperformed its larger V4-Pro flagship on agent-focused benchmarks through additional post-training rather than architectural changes. The model achieved 82.7 on Terminal-Bench 2.1, 54.4 on DeepSWE, and 70.3 on Toolathlon-Verified, though independent testing found lower scores of 79% on Terminal-Bench 2.1. The open-weight release under MIT license gives organizations direct control over deployment and integration with existing OpenAI-style APIs, reducing switching costs and infrastructure requirements.