TLDRocket
Sign in

Mistral OCR

Mistral AI

Mistral just launched an OCR API that reads PDFs, images, tables, math, and multiple languages at once. It beats Google, Azure, and GPT-4o on accuracy while processing 2,000 pages a minute.

Based on reporting by Mistral AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Mistral AI wants to solve a problem that sounds boring until you realize how much money is buried in it. Roughly 90% of organizational data sits locked inside documents — scanned contracts, scientific papers, lecture slides, engineering drawings — and most of it is functionally invisible to AI systems. Mistral OCR, launched this week, is the company's bet that turning those PDFs into structured, machine-readable text is worth building an entire product around.

The pitch isn't just "we read text." Mistral OCR pulls out embedded images, LaTeX equations, tables, and complex layouts, then hands back an interleaved markdown-style output that keeps everything in its original order. That matters a lot for anyone building retrieval systems, because a RAG pipeline that ignores charts and figures is only getting half the document. Mistral is already running this as the default document engine inside Le Chat, which tells you they trust it enough to put in front of millions of users without a fallback.

On the numbers, Mistral claims a clear lead. In internal text-only benchmarks measuring math, multilingual accuracy, scanned documents, and tables, Mistral OCR 2503 scored 94.89 overall, ahead of GPT-4o-2024-11-20 at 89.77, Gemini-2.0-Flash-001 at 88.69, and Azure OCR at 89.52. The gap widens on scanned documents (98.96 versus Gemini's 95.11) and tables (96.12 versus GPT-4o's 91.70). Language coverage looks strong too — Mandarin hit 97.11% fuzzy-match accuracy compared to roughly 91-92% for the nearest competitors, which is a meaningful jump for anyone dealing with non-Latin scripts at scale.

Speed is the other half of the argument. Mistral says the model processes up to 2,000 pages per minute on a single node, and prices the API at 1,000 pages per dollar, cheaper still with batch inference. That combination of speed and cost is aimed squarely at high-throughput use cases — research institutions digitizing journal archives, customer service teams indexing manuals, cultural heritage groups scanning historical records. Mistral is also pushing a

My take — AI-written commentary, not fact-checked reporting

, oh wait — the flashier feature here is treating documents themselves as prompts, letting users pull structured JSON straight out of a PDF and chain it into agent workflows without a separate parsing step. Availability is the usual Mistral playbook: open on la Plateforme today, headed for cloud and inference partners, and — notably — offered as a selective self-hosted deployment for organizations handling classified or sensitive material. That last option is clearly aimed at government and enterprise customers who won't send documents to anyone's cloud, and it's a distinguishing move against OCR offerings from Google and Microsoft that don't typically offer that kind of on-prem flexibility. most kind of unglamorous infrastructure play that ends up mattering more than the flashy demo, because unlocking a pile of PDFs is a real bottleneck, not a hypothetical one., "opinion":"OCR is not the topic anyone claims to plug",

Read more about this at: Mistral AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.