New coding and agent benchmarks using private, tool-enabled tasks show large drops in pass rates and limited consistency across repeated runs
Benchmark result Provisional 74% confidence first seen
Coverage reports new evaluations of AI coding agents on task benchmarks that include private, production-like codebases and full tool setups. Results show substantial failure rates on many attempts and a consistency gap when tasks are repeated, alongside efforts to improve repeat reliability through added analysis and guidelines.