[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
Latent Space ● Covered by 5 sources
DeepSeek shipped V4.1 Flash, a new open-weight model with text and vision. The twist: it’s built to cut active compute and KV-cache costs hard.
Based on reporting by Latent Space — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
DeepSeek is back with a model that looks like a modest version bump and behaves like a reset. V4.1 Flash is an open-weight flagship with text and image input, a 1M-token context, MIT licensing, and first-party US/API availability. Benchmarks are already framing it as a cheaper, more efficient successor to V4 Pro 0813, not just a small update.
The main story is the architecture. Independent benchmark reporting describes V4.1 Flash as a causal encoder-decoder design with 763B total parameters, but only 8B active on input and 16B active on output. That split is the point. DeepSeek is pushing more work into the right place at the right time, trying to keep inference lean while still handling long context and multimodal input.
That focus shows up in the efficiency numbers people are sharing. Artificial Analysis says the model scores 40 on its Intelligence Index, lands above the latest V4 Pro, and comes in at $0.30 per 1M input tokens and $1.20 per 1M output tokens, with cached input at $0.006 per 1M and an extra 50% off-peak discount. It also reports a 69% AutomationBench-AA score, a 1632 Elo on GDPval-AA v2, and 84% on AA-LCR v1.1. The catch is verbosity: AA says V4.1 Flash averages 89k tokens per task, which is a lot of talking for a model being sold on efficiency.
Even so, the cost picture is hard to ignore. AA estimates about $0.27 per Intelligence Index task, far below GLM-5.3 and Kimi K3, while Vals says it is the top open-weight model on its board at $0.30 per test. Baseten and Ollama moved quickly on support, and the local-running reports are what really made the launch land with systems people. One user claimed 200 TPS on four Max-Qs with 64GB of system RAM, then 300+ TPS on four RTX Pros, while another said it ran unexpectedly fast on a 128GB M5 Max with SSD streaming.
That is the DeepSeek pattern in a nutshell: less glamour, more machinery. The company keeps publishing architectures that make inference cheaper, then lets everyone else argue about whether the headline numbers are impressive enough. In this case, the hardware nerds seem to have answered first.
My take — AI-written commentary, not fact-checked reporting
This is the kind of release open-model advocates should want more often: fewer victory laps, more engineering that actually saves money. The industry loves benchmark theater, but the real flex is making a 763B model feel less like a whale and more like something an infra team can breathe around. That’s the part the closed labs hate most, and they’re right to be annoyed.
Read more about this at: Latent Space
Related stories
DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains
MarkTechPost · 1 month ago ·
51
DeepSeek-V4-Flash Outshines Pro, The Biggest GitHub Crawl Yet, Engineering System Prompts for Safer Code
The Batch ·
43