AI2 and Hugging Face introduce BenchMIRT, a method to audit what LLM benchmark scores measure at the level of individual prompts
Research publication Provisional 74% confidence first seen
Allen Institute for AI (AI2) and Hugging Face describe BenchMIRT, a technique for analyzing LLM benchmark results to estimate which underlying capabilities each prompt tests. The method is trained on outputs from 100 language models across 16 benchmarks and over 34,000 questions, and it identifies stable capability dimensions such as safety and general reasoning, showing how averaging benchmark items can obscure mixed signals.