Cohere Releases Embed 5: How It Compares to Voyage 4 Large, Gemini Embedding 2, and OpenAI
MarkTechPost Sana Hassan ● Covered by 2 sources
Cohere launched Embed 5, a new model family for search and RAG. The hook: Pro and Fast share one vector space, so you can index once and query cheaper.
Based on reporting by MarkTechPost, Sana Hassan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Cohere’s new Embed 5 family is built for enterprise search, retrieval-augmented generation, and agentic search flows. It comes in two tiers: Embed 5 Pro for maximum retrieval quality, and Embed 5 Fast for lower latency and cost on the live query path. Both accept text, images, and mixed text-plus-image input, cover more than 100 languages, and handle up to 128K tokens.
The most interesting design choice is that Pro and Fast use the same embedding space. That means one model can build the index and the other can query it, without forcing a separate pipeline. Cohere says its recommended setup is to index with Pro and query with Fast, as long as both sides use the same output dimension. Both tiers are generally available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker, with private VPC or on-prem serving through vLLM.
Cohere is also pushing flexibility on the vector format itself. The API model IDs are embed-v5.0-pro and embed-v5.0-fast. Both can output 2048, 1536, 1024, 768, 512, or 256 dimensions, with 2048 as the default, and embeddings can come back as float, int8, or binary. Pricing splits cleanly too: Pro is $0.12 per 1M text tokens, Fast is $0.08, and image input costs $0.40 per 1M tokens on both.
That setup is clearly aimed at real-world retrieval systems where storage and latency matter. Cohere says a 2048-dim float32 vector needs 8 KB, while a 1024-dim int8 vector needs 1 KB and a 256-dim binary vector needs 32 bytes. Across 100M chunks, raw storage falls from about 819 GB to 3.2 GB. Cohere recommends 1024-dim int8 as the default because it keeps most of the quality while cutting the footprint hard.
On benchmarks, Pro does well. On ViDoRe V3 it averages 85.8, which Cohere says is 8.8 points better than Embed 4. Fast averages 84.5. Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2, and OpenAI text-embedding-3-large scores 75.5 on that same benchmark. But the picture is messier outside English: Cohere says Pro leads the European-language average, while Gemini Embedding 2 beats it on 9 of 10 additional languages in Cohere’s own table. Cohere also notes that most of these figures use a new metric, RCP-nDCG@10, which reorders a fixed candidate set and looks more like reranking than first-stage retrieval.
My take — AI-written commentary, not fact-checked reporting
This is the sane way to ship embeddings: one space, two cost tiers, and fewer excuses for vendor lock-in theater. The catch is the usual one—vendor benchmarks are nice, but until independent runs land, they’re still vendor benchmarks. Still, Cohere at least seems to understand that retrieval systems live or die on boring things like latency, storage, and not making teams rebuild the whole stack every time a model changes.
Read more about this at: MarkTechPost