TLDRocket
Sign in

Ground truth is a process, not a dataset

Amazon Science

Amazon's AGI group developed a new evaluation method for AI-generated research reports, discovering that traditional static benchmarks fail when assessing complex factuality claims that require cross-document synthesis. Expert evaluators achieved only 60.8% accuracy on known answers using standard labeling, but improved to 90.9% accuracy when placed in an auditing role that compared competing evidence. The audit-then-score protocol treats ground truth as a dynamic process where models can challenge benchmark answers with evidence, fundamentally changing how AI evaluation works and enabling DeepFact-Eval to reach 83.4% accuracy on their benchmark.

Why it matters

Automatically fact-checking long, AI-generated research reports poses new challenges — including benchmarking.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.