TLDRocket
Sign in

DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed

The New Stack Jessica Wachtel Covered by 3 sources

DeepSeek just added image input to its cheap V4 Flash model. It matched Gemini on accuracy, but Gemini was faster and DeepSeek was cheaper.

Based on reporting by The New Stack, Jessica Wachtel — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

DeepSeek rolled out V4 Flash Vision Exp on August 21, then pushed it to API gateways like OpenRouter on August 27. It’s the company’s first model that can take images, so charts, screenshots, and photo documents can be handled alongside text. The pitch is simple: keep the budget pricing of V4 Flash and add vision on top.

The pricing is still $0.22 per million input tokens and $0.66 per million output tokens, though weekday peak hours double that. Google’s Gemini 3.7 Flash, out on August 13, is the obvious rival here. On OpenRouter it runs at $0.75 and $3.75 per million, which makes it the pricier option before you even start measuring speed.

To compare them, the test used three image-heavy tasks that look a lot like real office work: reading a stacked bar chart with a second y-axis, auditing a flawed invoice, and tracing the root cause in production logs. Both models cleared every trap. They identified the same quarter on the chart, the same growing revenue segment, the same full-year revenue, and they both caught the invoice errors and the log-file decoy.

Where they split was in how they got there. DeepSeek sometimes took a long time, especially on the invoice test, where it spent 30.5 seconds and produced 3,467 completion tokens for a short answer. Gemini answered that same prompt in 7.9 seconds with 944 completion tokens. On the logs test, DeepSeek took 11.9 seconds and Gemini 7.6 seconds. DeepSeek also counted images at about 500 prompt tokens each, while Gemini used about 1,150, so the two systems clearly treat image input very differently.

Across all three tests, the score was the same: 9/9 accuracy for both models. The real difference was the bill and the wait. DeepSeek’s total came to $0.0039, about a third of Gemini’s $0.0122, while Gemini averaged 7.2 seconds and DeepSeek 16.8 seconds. That leaves a pretty clean tradeoff. If the job is batch invoice work and nobody is staring at a loading spinner, DeepSeek looks good. If someone is waiting on the answer, Gemini wins on pace.

My take — AI-written commentary, not fact-checked reporting

This is the old AI bargain in a cleaner outfit: cheap, slow, or fast, pricier. DeepSeek’s vision debut looks solid, but weekday peak pricing and wildly uneven response times are not exactly confidence-inspiring for production use. The industry keeps pretending every model is “best” until the stopwatch and the invoice show up.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.