Automating Eval Design and Hillclimbing with Claude
claude.dev Blog
Claude Code’s claude-api skill added guided workflows for building evals and improving apps via /claude-api build-eval and /claude-api hillclimb. The hillclimb workflow warns when a baseline score is around 95% or higher. It now produces structured eval cases, grader checks, and an iterated patch process that splits data into train and test to reduce overfitting before keeping changes.
Why it matters
The post says good evaluations should mirror production, preserve headroom, reward stronger models and more thinking, and keep low run-to-run variance. It introduces /claude-api build-eval and /claude-api hillclimb to build reviewed test sets and iteratively improve prompts, skills, model settings, or harness code.
Related stories
Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
MarkTechPost · 2 weeks ago ·
34
AlignEval: Building an App to Make Evals Easy, Fun, and Automated
Eugene Yan · 1 year ago ·
29
Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
MarkTechPost · 1 month ago ·
23