TLDRocket
Sign in

Data Machina #254

Substack

Princeton Language & Intelligence released SWE-bench, a benchmark for evaluating AI coding agents on their ability to fix real GitHub repository issues, revealing that current AI agents perform poorly on the task. The leading model, Amazon Q Developer Agent, successfully solved only 13.8% of 2294 tasks, while the open-source OpenDevin achieved the highest benchmark score at 21%. These results demonstrate that despite multiple competing approaches including Devin, Devika, and GPT-Engineer, AI coding agents remain far from ready for production-scale legacy code migration and autonomous software engineering work.

Why it matters

State of AI Coding Agents. SWE-Agent. Amazon Q. Devin. OpenDevin. Devika. Blackbox AI. GPT-Engineer. ChatDev. KHOJ Personal AI Agents. Perplexica. CogVLM2. World Models.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.