TLDRocket
Sign in

Better Models: Worse Tools

Armin Ronacher

Newer Anthropic Claude models (Opus 4.8 and Sonnet 5) sometimes fail to properly call external tools by inventing extra fields in the API schema, whereas older models did not exhibit this problem. The failure rate reaches approximately 20% in some agentic contexts and can be reduced to zero by enabling strict tool invocation mode. This regression appears to result from post-training on Claude Code's forgiving tool harness, which silently repairs malformed calls and accepts parameter aliases, causing newer models to learn looser compliance with unfamiliar tool schemas.

Why it matters

Newer models get better at solving tasks while getting worse at faithfully emitting alternative tool schemas, requiring stronger guarantees in the harness.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.