TLDRocket
Sign in

Claude Opus 5.5 wants to finish your coding tasks, not just start them

The New Stack Adrian Bridgwater ● Covered by 8 sources

Anthropic just pushed Claude Opus 5.5 toward finishing whole coding jobs, not just spitting out snippets. That matters because it’s cheaper than Opus 5 and aimed at bigger, messier dev work.

Based on reporting by The New Stack, Adrian Bridgwater — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic is trying to move Claude from a code helper into something closer to a full project worker. On Tuesday, the company launched Claude Opus 5.5, the first model in a new Claude 5.5 family, with a pitch that reaches across the software development process: design specs, debugging, code generation, and testing.

The headline for developers is that it’s cheaper and more capable on long jobs. Anthropic says Opus 5.5 matches Claude Fable 5.1 on “most work” tasks while costing about 40% less to run than Opus 5. In its own launch material, the company says the model is especially strong on sprawling tasks such as codebase-wide migrations and audits.

Anthropic also leaned hard on examples. An early tester said Opus 5.5 audited and fixed a 200,000-line codebase in under three hours, while Opus 5 needed more than 20 hours and used 2.5x as many tokens. In another internal test, the company says Opus 5.5 translated HAProxy from C to Rust, passed nearly all of HAProxy’s regression tests, finished in 9.5 hours, and cost 51% less than Fable 5.1.

There’s a safety story here too, and Anthropic is clearly using it to justify the model’s broader reach. The company says Opus 5.5 was built with alignment testing, outside pre-release evaluation, and extra safeguards around cybersecurity and biology. It says the model is the strongest performer it has tested on its most comprehensive alignment test, with improvements in behaviors tied to recent cybersecurity incidents, including biased reasoning and attempts to escape a sandbox.

But the bigger shift may be how developers are supposed to use it. Anthropic and several testers describe a workflow where the terminal is no longer just a place to paste code, but part of the model’s workspace. That means the agent can run commands, inspect failures, edit files, and check its own work instead of handing off a half-finished mess.

Pricing pushes the same direction. Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens, with cache read prices cut by 60% for token-billed usage. Output is also said to be more than 30% faster than Opus 5. The message is simple: Anthropic wants Claude to finish the ticket, not just warm it up.

My take — AI-written commentary, not fact-checked reporting

This is the real AI software story now: less autocomplete theater, more governed work inside real systems. The catch is that “finished” is not the same as “correct,” and companies pretending otherwise are just outsourcing their bugs with extra confidence. The sensible move is to treat these models like fast junior engineers with a dangerous sense of closure.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.