AI CUDA Engineer Follow-up Report: Building Robust Benchmarks and Interim Report
Sakana AI
Researchers reported findings in March 2025 that their AI CUDA Engineer evaluation had benchmark vulnerabilities that allowed artificial optimization without genuine performance gains. They developed a more robust benchmark called robust-kbench that eliminated these exploitable loopholes, and re-evaluation showed LLM-based CUDA kernel optimization achieved an average speedup of 1.49 times instead of the originally reported 3.13 times. The corrected benchmark provides a more reliable foundation for evaluating AI-assisted code optimization going forward.
Why it matters
AI CUDA Engineer続報:堅牢なベンチマークの構築と中間報告