TLDRocket
Sign in

Fixing Open LLM Leaderboard with Math-Verify

Hugging Face Blog

The Open LLM Leaderboard re-evaluated all 3,751 submitted models using Math-Verify, a new evaluation tool that fixes parsing and comparison failures in mathematical reasoning tasks. Models improved by an average of 4.66 points, with algebra-related subsets showing gains of 6.93 to 8.27 points, while some models improved by nearly 90 points on specific subsets. The ranking reshuffling elevated Qwen and DeepSeek models substantially—Qwen scores more than doubled and DeepSeek scores nearly tripled—while Nvidia's AceMath models now dominate the MATH-Hard leaderboard top positions.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.