TLDRocket
Sign in

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

MarkTechPost Asif Razzaq

Perplexity dropped WANDR, a new benchmark that tests AI research agents on building huge, evidence-backed lists, not just single answers. Turns out even the best agents only nail about a third of it.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Research agents are already doing real jobs today: scoping competitors, running due diligence, chasing down candidates. But the tests we use to grade them mostly check one clean answer, which is a strange way to evaluate a tool whose whole job is building giant lists of verified facts. Perplexity's new WANDR benchmark tries to fix that mismatch, and the results it published alongside the launch are a useful reality check on how far agentic search actually is from professional-grade reliability.

WANDR stands for Wide ANd Deep Research, and it's the companion to Perplexity's earlier DRACO benchmark. Where DRACO judges long-form reports, WANDR judges collections: 500 tasks built from a hierarchy like company(70) -> employee(1) -> url(1), meaning find 70 qualifying companies, get a specific hire at each one, and back it up with a source page. One example task asks for at least 70 US companies with a CEO or CFO appointment announced between March and April 2026, each needing an authoritative announcement page plus a listing page, for 140 total records. Across all 500 tasks, the benchmark demands 170,495 source-backed records, and grading isn't against a fixed answer key. A grader actually re-fetches every cited URL during evaluation and checks whether the excerpt genuinely supports the claim.

The scoreboard is humbling. Perplexity's own Search as Code system topped the field with a soft F1 of 0.363 and a hard F1 (meaning the entire branch of evidence checks out) of just 0.133, running about $5.20 and 15 minutes per task. Anthropic came closest on quality but burned more time, money, and tokens to get there. Everyone else, including OpenAI and Exa's setups, was faster and cheaper but noticeably worse. Push Perplexity's system to its priciest

My take — AI-written commentary, not fact-checked reporting

placeholder

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.