TLDRocket
Sign in

Introducing the LiveCodeBench Leaderboard - Holistic and Contamination-Free Evaluation of Code LLMs

Hugging Face Blog

Researchers from UC Berkeley, MIT, and Cornell released LiveCodeBench, a new leaderboard for evaluating code-generation capabilities of large language models across four tasks: code generation, self-repair, code execution, and test output prediction. The benchmark collects problems from LeetCode, AtCoder, and CodeForces with annotated release dates, enabling evaluation on problems released after a model's training cutoff to prevent contamination. GPT-4-Turbo performs best on most scenarios, while Claude-3-Opus excels at test output prediction and Mistral-Large shows stronger performance on natural language reasoning tasks.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.