Introducing SimpleQA
OpenAI 1 year ago 28
OpenAI released SimpleQA, a benchmark designed to measure how accurately language models answer short factual questions. The benchmark evaluates models on their ability to provide correct answers to straightforward queries that seek specific facts. This allows researchers to systematically test and compare language models' factual accuracy across different architectures and training approaches.