TLDRocket
Sign in

Separating signal from noise in coding evaluations

TLDR Dev Covered by 2 sources

Anthropic audited SWE-Bench Pro and found about 30% of tasks were broken due to overly strict tests and unclear specifications. The audit discovered issues that made the benchmark unsuitable for reliable evaluation, causing Anthropic to withdraw its recommendation for using SWE-Bench Pro. This finding prompted the community to reconsider which benchmarks should be used to measure the actual capabilities of software engineering AI systems.

Why it matters

A recent audit by Anthropic found that approximately 30% of tasks in SWE-Bench Pro are broken due to issues such as overly strict tests, underspecified prompts, and misleading instructions, leading to a retracted recommendation to transition to the benchmark.

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.