AI coding agents benchmarking (token use per solved task)
Reported model scores on AI coding agents benchmarking (token use per solved task), best parseable score first. Each row keeps its verbatim score, test conditions and provenance, and links to the source coverage it was extracted from. Benchmark profile →
| Model | Score | Conditions | Provenance | Measured | Source |
|---|---|---|---|---|---|
| OpenClaw | 292,000 tokens per solved task | — | independent | — | coverage → |
| Aider (architect mode) | about 3,500 tokens per solved task | architect mode | independent | — | coverage → |
Scores are only comparable within one benchmark under matching conditions — results under different test setups, and scores from other benchmarks, are not directly comparable.