Introducing SimpleQA
OpenAI Blog
OpenAI released SimpleQA, a benchmark designed to measure how accurately language models answer short factual questions. The benchmark evaluates models on their ability to provide correct answers to straightforward queries that seek specific facts. This allows researchers to systematically test and compare language models' factual accuracy across different architectures and training approaches.
Why it matters
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.