TLDRocket
Sign in

Agent Benchmarking

21 summarised stories about Agent Benchmarking, each linking back to the original source. Browse all topics →

Sunday, 19 July 2026

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

MarkTechPost 2 days ago

Perplexity released WANDR, an open benchmark with 500 tasks that evaluates research agents on their ability to discover large collections of entities and support claims with evidence. The benchmark requires 170,495 source-backed records across all tasks, with Perplexity's Search as Code system achieving a soft F1 score of 0.363 while other systems score significantly lower. The results show that discovery and extracting complete evidence from retrieved pages remain the primary challenges for current research agents.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.