TLDRocket
Sign in

Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

MarkTechPost Asif Razzaq

Perplexity released two new embedding models for text, images, and PDF pages. The small one is meant for edge devices; the bigger one hit 92.4% on MADQA.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Perplexity has put out pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models that share one embedding space and can handle text, images, and rendered PDF pages. The release comes in two sizes: a 0.6B model aimed at fast, cheap queries, and a 9B model aimed at better quality at index time.

The smaller model is the one that sounds most practical in the real world. Perplexity says it uses about 240M active parameters for text and 340M for images, and that it can run on a laptop or edge device. The 9B model is meant for datacenter or high-memory GPU setups. Both checkpoints are on Hugging Face under the MIT license, with commercial use allowed, and Perplexity says a hosted API endpoint is planned but not live.

The headline score is 92.4% on MADQA for the 9B model, which Perplexity says is best among the models it tested. On the same benchmark, the 0.6B model scores 90.1%. The same split shows up elsewhere too: the 9B model leads on domain-specific text with 81.3% nDCG@10, while the 0.6B model lands at 78.0%. On Q2D-Web, the two models score 74.8% and 73.6% recall@1000.

The trade-offs are easy to see. Each token gets its own 128-dimensional vector, which keeps the space narrow but means index size grows with document length. That helps explain why the models are good fits for visual document search over PDFs, slides, and scanned reports, but not a clean win everywhere. Perplexity says a 9B index queried by the 0.6B model can recover about half the quality gap on text at the cost of 0.6B queries, and that setup scored 63.5% on ViDoRe v3 image retrieval. Still, image search remains the rough edge: Tencent’s EVIE scores higher on ViDoRe v3 image retrieval, and Gemini Embedding 2 beats Perplexity on MIRACL-Vision and PPLX-Q2I.

Perplexity also says the models were distilled from an 18B teacher with LEAF-style token-level training, which is what creates the shared space. One catch: a single input can’t mix text and images, and the scores are self-reported because the technical report is not out yet.

My take — AI-written commentary, not fact-checked reporting

This is the kind of release that makes more sense than half the “AI for everything” theater floating around. A small model for local queries and a bigger one for index quality is a clean split, even if the storage story is a bit hungry. The real tell is that the useful part here is retrieval over messy documents, not another shiny chatbot demo.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.