TLDRocket
Sign in

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

MarkTechPost Asif Razzaq

Perplexity released WANDR, an open benchmark with 500 tasks that evaluates research agents on their ability to discover large collections of entities and support claims with evidence. The benchmark requires 170,495 source-backed records across all tasks, with Perplexity's Search as Code system achieving a soft F1 score of 0.363 while other systems score significantly lower. The results show that discovery and extracting complete evidence from retrieved pages remain the primary challenges for current research agents.

Why it matters

Perplexity's WANDR is an open benchmark and evaluation harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each one with cited, re-verifiable evidence. Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1. The post Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep appeared first on MarkTechPost.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.