Agents on Rails benchmark leaderboard updated with rerun results for Anthropic Claude Fable 5.1 and Z.ai GLM 5.3 Flash
Benchmark result Provisional 74% confidence first seen
The Agents on Rails benchmark leaderboard was updated after rerunning and reposting results for Anthropic’s Claude Fable 5.1 and Z.ai’s GLM 5.3 Flash. Coverage reports changes in solve rates, median run times, and total costs, with Fable 5.1 listed near the top on multiple price and security-related performance marks, while an independent budget-focused comparison found a smaller advantage than the benchmark’s headline claim.