The Hallucinations Leaderboard, an Open Effort to Measure Hallucinations in Large Language Models
Hugging Face Blog
Researchers created the Hallucinations Leaderboard, an open platform that evaluates large language models on their tendency to generate false or inaccurate content through a series of standardized benchmarks. The leaderboard tests models across 12 categories including factual question-answering, summarization, reading comprehension, and hallucination detection tasks, with scores normalized on a 0–1 scale. The platform enables researchers and developers to compare model reliability and identify which models are most prone to producing misinformation or contradicting user instructions.