TLDRocket
Sign in

AI Evaluation

13 summarised stories about AI Evaluation, each linking back to the original source. Browse all topics →

+ Follow this topic

Wednesday, 5 August 2026

AI agents can't yet do open-ended AI research

AI as Normal Technology 3 weeks ago 48

Researchers tested frontier AI agents on open-ended AI research tasks by having them answer unpublished research questions, and both agents' papers were rejected by the original authors. The agents had $1,000+ in API credits and six days to work but spent less than 50% of their budget and failed to show creativity, backtracking, or effective judgment when facing setbacks. The findings suggest that recursive self-improvement through AI research remains limited to narrow, verifiable tasks, and open-ended research remains a bottleneck that could slow explosive AI progress.

What Actually Keeps an AI Benchmark Useful? Scale

StackSweep 3 weeks ago 39

Nearly half of 60 widely used language model benchmarks have become saturated, meaning top models score within statistical noise of each other and the benchmark can no longer rank them. Of the 60 benchmarks analyzed, 29 show high or very high saturation (Saturation Index ≥0.7), with saturation climbing from a mean of 0.51 for benchmarks under 24 months old to 0.60 for those over 60 months old. The study found that assumed safeguards like private test sets, harder output formats, and multilingual scope provide no meaningful protection once benchmark age is controlled for, leaving only test set scale and expert curation as effective strategies, and recommending explicit retirement criteria rather than letting benchmarks accumulate citations indefinitely.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.