TLDRocket
Sign in

Agent Evaluation

20 summarised stories about Agent Evaluation, each linking back to the original source. Browse all topics →

Wednesday, 18 February 2026

IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST

Hugging Face Blog 5 months ago

IBM and UC Berkeley created MAST, a diagnostic framework that categorizes why AI agents fail in IT automation tasks rather than just reporting success rates, and applied it to analyze 310 execution traces across three language models. Gemini-3-Flash averaged 2.6 failure modes per failed trace while GPT-OSS-120B averaged 5.3, showing that smaller open-source models suffer from cascading failures that compound over time. The analysis revealed that incorrect verification (agents declaring success without checking results) is the strongest failure predictor, enabling developers to deploy targeted fixes like external verification gates instead of blind prompt engineering.

Introducing EVMbench

OpenAI Blog 5 months ago

OpenAI and Paradigm released EVMbench, a benchmark designed to measure how well AI agents can identify, fix, and exploit critical vulnerabilities in smart contracts. The benchmark includes high-severity flaws across multiple smart contract categories to test AI performance on security tasks. This creates a standardised way to evaluate whether AI systems can assist with smart contract security analysis and auditing.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.