TLDRocket
Sign in

Data Machina #254

Data Machina

Princeton Language & Intelligence released SWE-bench, a benchmark for evaluating AI coding agents on their ability to fix real GitHub repository issues, revealing that current AI agents perform poorly on the task. The leading model, Amazon Q Developer Agent, successfully solved only 13.8% of 2294 tasks, while the open-source OpenDevin achieved the highest benchmark score at 21%. These results demonstrate that despite multiple competing approaches including Devin, Devika, and GPT-Engineer, AI coding agents remain far from ready for production-scale legacy code migration and autonomous software engineering work.

Why it matters

State of AI Coding Agents. SWE-Agent. Amazon Q. Devin. OpenDevin. Devika. Blackbox AI. GPT-Engineer. ChatDev. KHOJ Personal AI Agents. Perplexica. CogVLM2. World Models.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.