TLDRocket
Sign in

Agent Evaluation

39 summarised stories about Agent Evaluation, each linking back to the original source. Browse all topics →

+ Follow this topic

Monday, 31 August 2026

Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour

MarkTechPost 3 days ago 1

Keenable AI released NEEDLE, an open-source live web search benchmark that avoids leakage by regenerating its query sets from fresh public sources while preventing engines from fetching or re-ranking results. The benchmark regenerates news queries hourly (from RSS feeds and Google Trends) and scores 15 search APIs against an “ultimate” pooled oracle ceiling, with evidence clipped to 2,000 characters. Results are now reported against a moving, field-wide achievable ceiling rather than a static test set, changing evaluation to separate retrieval failures from ranking failures across the market.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.