TLDRocket
1 August 2026
DeepSeek's V4-Flash 0731 update reshuffled the economics of AI inference in a single afternoon. The post-training refinement lifted the model's Terminal-Bench score from 56.9 to 82.7 points—a 45 percent jump without touching its 284 billion parameters—and the company immediately released weights under MIT, flooding the market with a capable open model priced at $0.14 per million input tokens. That's undercutting even last month's aggressive pricing, but the real shift is subtler: developers are no longer hunting for better base models. They're engineering routing systems and inference harnesses to squeeze more from what they already have. DeepSeek's move signals that the marginal value of raw capability gains has collapsed; what matters now is latency, cost-per-task, and the plumbing that decides which model handles which query. The price floor—and the fierce velocity of these updates—leaves proprietary vendors scrambling to justify premium pricing on closed systems. Open-weights inference has stopped being a scrappy alternative and become the default assumption, forcing the conversation upstream to orchestration and application design rather than model weights themselves.
Read the full briefing →