TLDRocket
Sign in

DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL

Together AI

Agentica and Together AI released DeepSWE-Preview, a 32B parameter coding agent trained entirely with reinforcement learning that achieves 59% accuracy on SWE-Bench-Verified with test-time scaling. The model was trained on 4,500 real-world software engineering tasks over six days using 64 H100 GPUs, reaching 42.2% Pass@1 performance as the best open-weight coding agent. The team open-sourced the dataset, training code, and evaluation logs, along with improvements to the GRPO algorithm including techniques like compact filtering and length normalization for stable multi-turn agent training.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.