Agents on Rails
Reported model scores on Agents on Rails, best parseable score first. Each row keeps its verbatim score, test conditions and provenance, and links to the source coverage it was extracted from. Benchmark profile →
| Model | Score | Conditions | Provenance | Measured | Source |
|---|---|---|---|---|---|
| Claude Fable 5.1 | 92% (58 of 63 runs) | $75 total price for all 63 runs; 5.4 minutes median per run (rerun leaderboard results) | independent | — | coverage → |
| Z.ai GLM 5.3 Flash | 83% (52 of 63 runs) | $3.31 total cost for all 63 runs (rerun leaderboard results) | independent | — | coverage → |
Scores are only comparable within one benchmark under matching conditions — results under different test setups, and scores from other benchmarks, are not directly comparable.