Researchers from IBM, Hugging Face, and other institutions launched EveryEvalEver, a standardized database and JSON format for documenting AI model benchmark results to address inconsistencies in how evaluation scores are reported across the industry. The platform currently aggregates 22,000 model results across 2,200 benchmarks translated from 31 different evaluation formats, revealing that identical evaluations can produce scores varying by as much as 20 percentage points due to different harnesses and undocumented settings. By establishing a common language for benchmarking metadata and making prior results reusable, the project aims to reduce the $370,000 cost of reproducing aggregated evaluations and improve transparency in model comparisons.
A researcher tested whether AI labs optimize their models for Simon Willison's famous "pelican riding a bicycle" benchmark by generating 1,008 SVG images across 48 animal-vehicle combinations from seven models and scoring them with an LLM judge. The pelican-on-bicycle combination ranked 42nd of 48 in overall quality, and statistical analysis found no significant per-lab boost for pelicans, bicycles, or their combination after adjusting for difficulty. The results suggest AI labs are not noticeably optimizing for this benchmark, though a small non-significant effect appeared in one model.
Multiple open-source speech recognition models now compete at similar accuracy levels, with Cohere's Transcribe (5.42% WER), IBM's Granite Speech 4.1 (5.33%), and others within one percentage point of each other. The Open ASR Leaderboard rankings are unreliable because models are evaluated on different test sets—excluding easier benchmarks like TED-LIUM artificially inflates some scores. Model selection now depends on license type, language support, streaming capability, and cost per audio-hour rather than benchmark rank, making this a procurement decision rather than a research one.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.