New benchmarks and benchmark results were published to evaluate AI agents’ end-to-end commercial performance and their ability to build other agents
Benchmark result Provisional 48% confidence first seen
Alibaba’s Accio team released CommerceAgentBench, an open-source benchmark that scores AI agents based on end-to-end outcomes in e-commerce tasks rather than only model reasoning. Separate coverage also reported results on Hyper-τ-bench, which tested whether AI developer agents can build customer-service agents from business materials, with top systems still achieving low success rates.