TLDRocket
Sign in

Model Evaluation

56 summarised stories about Model Evaluation, each linking back to the original source. Browse all topics →

+ Follow this topic

Friday, 1 December 2023

Open LLM Leaderboard: DROP deep dive

Hugging Face 2 years ago 25

The Open LLM Leaderboard discovered critical flaws in its DROP benchmark implementation, where most models scored below 10 out of 100 on the f1-score metric despite appearing capable on other benchmarks. The investigation identified two main issues: the normalization step failed when numbers were followed by non-space whitespace characters, and using a period as the stop token prevented models from completing floating-point answers and generated extraneous text. The benchmark has been removed from the leaderboard pending development of a corrected evaluation implementation, as fixing the issues would require rerunning over 50 percent of test cases.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.