Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Google Research
Knowledge profiling found that many factual mistakes in frontier LLMs are recall failures (encoded facts not reliably accessible) rather than encoding failures (facts not encoded). 95–98% of facts were encoded in Gemini-3-Pro and GPT-5, yet direct recall failed on 26–34% of facts. The implication is a shift toward interventions that improve knowledge utilization (including thinking/recovery) instead of primarily scaling to add missing knowledge.
Why it matters
Generative AI