TLDRocket
Sign in

FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

Google DeepMind

Researchers released the FACTS Benchmark Suite, a new evaluation system comprising 3,513 examples designed to measure how accurately large language models answer factual questions across four categories: parametric knowledge, web search, multimodal reasoning, and grounded responses. Gemini 3 Pro achieved the highest overall FACTS Score of 68.8%, with error rates 55% lower than Gemini 2.5 Pro on search-based questions and 35% lower on parametric questions. The benchmark suite, now hosted on Kaggle, provides a public leaderboard to track LLM factuality improvements over time.

Why it matters

Systematically evaluating the factuality of large language models with the FACTS Benchmark Suite.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.