TLDRocket
Sign in

Cohere’s faster query model barely dents retrieval quality in its tests

The New Stack Amanda Caswell

Cohere launched Embed 5, letting teams index with Pro and query with cheaper Fast. Its tests say the speed boost costs very little retrieval quality.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Cohere’s new Embed 5 gives teams a practical split: build the index with Embed 5 Pro, then run queries with the faster, cheaper Embed 5 Fast. That means one vector store, not two. For RAG systems and agent workflows, where the same data gets hit over and over, the company is betting that latency matters more on the query side than on the indexing side.

The tradeoff is real, but Cohere’s numbers make it look small. Across 40 datasets spanning text, images, fused documents, and parsed documents, Fast queries against a Pro index scored 98.4 against a Pro-to-Pro baseline of 100. Using Fast for both indexing and queries brought that down to 96.6. Cohere says none of the individual datasets showed a major drop when Pro and Fast were used together.

The appeal is not just quality. Pro and Fast share an embedding space, so teams can switch between them without re-embedding the corpus. They also produce compatible vectors at the same dimensions, and Cohere says they can be mixed with Matryoshka truncation or int8 quantization. Pricing tilts the same way: Pro is $0.12 per million tokens, Fast is $0.08, and Cohere says Fast averaged 2.4 times the document throughput in its tests.

Cohere is also pushing the storage angle hard. Both models support six vector dimensions from 256 to 2,048, with float32, int8 and binary formats. In Cohere’s figures, a 2,048-dimensional float32 vector works out to 8 KB, or roughly 819 GB for 100 million chunks. A 1,024-dimensional int8 vector drops that to about 102 GB, while a 256-dimensional binary vector brings it down to roughly 3.2 GB. The company recommends 1,024-dimensional int8 for most deployments, with binary left for earlier retrieval stages before reranking.

There’s more under the hood, too. Embed 5 handles text, images, and fused text-image inputs across more than 100 languages, with a 128K-token context window. Cohere says Pro and Fast both beat Google’s Gemini Embedding 2 on some of its fused and parsed-document tests, though the multilingual results are less clean and Pro trails Gemini Embedding 2 on nine of the tests. The company also warns that some of these scores use RCP-nDCG@10, which is not the same thing as first-stage retrieval from a full corpus, so the leaderboard glow needs a little squinting.

The bigger story is architectural. Cohere wants indexing and serving to become separate decisions: optimize quality when the corpus goes in, optimize throughput when users start asking questions. That is a sensible line to draw, and also a quietly brutal reminder that most “AI infrastructure” wins are really storage and latency arguments wearing a nicer shirt.

My take — AI-written commentary, not fact-checked reporting

This is the rare AI launch that sounds like infrastructure, not marketing confetti. Cohere is basically saying the expensive model can stay on the back end while the cheaper one does the busywork, which is exactly the kind of split serious teams should want. The industry keeps acting like every problem needs a grander model; sometimes it just needs fewer expensive tokens and less fanfare.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.