TLDRocket
Sign in

Featuring Every Eval Ever Results on Hugging Face Model Pages

Hugging Face Blog

Every Eval Ever and Hugging Face Community Evals have integrated their evaluation result systems to enable cross-posting and linking of benchmark scores across platforms. The combined datastore now contains approximately 229,000 evaluation results across 22,000 models and 2,200 benchmarks, drawn from 31 different reporting formats. Users can now submit evaluation results to both platforms simultaneously using a converter tool, with results appearing on model pages and linking back to full standardized records for reproducibility and interpretation.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.