TLDRocket
Sign in

Introducing the Open Chain of Thought Leaderboard

Hugging Face Blog

A new leaderboard launched to measure how effectively language models generate step-by-step reasoning traces, rather than scoring absolute accuracy on multiple-choice reasoning tasks. The evaluation uses four benchmarks from AGIEval (LogiQA and LSAT subsets) and measures accuracy gain as the difference between performance with and without chain-of-thought prompting across six generation regimes combining two prompting strategies and three decoding methods. Early results with 30 models show that smaller 7-billion-parameter models sometimes outperform larger ones at chain-of-thought reasoning, instruction-tuning improves both baseline and marginal gains, and no single prompting approach works best across all models and tasks.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.