TLDRocket
Sign in

Fixing Open LLM Leaderboard with Math-Verify

Hugging Face

The Open LLM Leaderboard re-evaluated all 3,751 submitted models using Math-Verify, a new evaluation tool that fixes parsing and comparison failures in mathematical reasoning tasks. Models improved by an average of 4.66 points, with algebra-related subsets showing gains of 6.93 to 8.27 points, while some models improved by nearly 90 points on specific subsets. The ranking reshuffling elevated Qwen and DeepSeek models substantially—Qwen scores more than doubled and DeepSeek scores nearly tripled—while Nvidia's AceMath models now dominate the MATH-Hard leaderboard top positions.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.