TLDRocket
Sign in

Automating Eval Design and Hillclimbing with Claude

claude.dev Blog

Claude Code’s claude-api skill added guided workflows for building evals and improving apps via /claude-api build-eval and /claude-api hillclimb. The hillclimb workflow warns when a baseline score is around 95% or higher. It now produces structured eval cases, grader checks, and an iterated patch process that splits data into train and test to reduce overfitting before keeping changes.

Why it matters

The post says good evaluations should mirror production, preserve headroom, reward stronger models and more thinking, and keep low run-to-run variance. It introduces /claude-api build-eval and /claude-api hillclimb to build reviewed test sets and iteratively improve prompts, skills, model settings, or harness code.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.