AppWorld
Reported model scores on AppWorld, best parseable score first. Each row keeps its verbatim score, test conditions and provenance, and links to the source coverage it was extracted from. Benchmark profile →
| Model | Score | Conditions | Provenance | Measured | Source |
|---|---|---|---|---|---|
| GPT-4.1 (ReAct agent) | 81.0% | after introducing Consistency Analyzer and consistency guidelines (Mean@5) | vendor-reported | — | coverage → |
| GPT-4.1 (ReAct agent) | 77.4% | average completion rate across runs | vendor-reported | — | coverage → |
| GPT-4.1 (ReAct agent) | 69.0% | after introducing Consistency Analyzer and consistency guidelines (Pass^5) | vendor-reported | — | coverage → |
| GPT-4.1 (ReAct agent) | 53.0% | tasks completed in all 5 repeated runs (Pass^5) | vendor-reported | — | coverage → |
Scores are only comparable within one benchmark under matching conditions — results under different test setups, and scores from other benchmarks, are not directly comparable.