Anthropic adds build-eval and hillclimb commands to claude-api skill for automated evals and optimization
Key point
The new commands automate eval design and iterative improvements, achieving up to 98.9% accuracy at one-fifth the cost in internal benchmarks.
Details
Anthropic has added build-eval and hillclimb commands to the claude-api skill, enabling automated evaluation design and iterative optimization of prompts, skills, or harness code. The tools guide users through creating robust eval sets from production data and systematically improving performance while guarding against overfitting and reward hacking.
Eval Design Principles
The build-eval command generates evaluation sets by sampling from production transcripts, bug reports, and codebase data. It supports both programmatic verification for constrained outputs and LLM-as-judge for open-ended tasks, requiring rubrics based on checkable claims rather than simple scales. The workflow emphasizes low run-to-run variance and ensures evals retain headroom by avoiding tasks that are impossible or ambiguous.
Hillclimbing Workflow
The hillclimb command iteratively improves performance or reduces cost by analyzing train transcripts and proposing single patches. It maintains a train/test split to detect overfitting, rolling back changes if test scores remain flat while train scores rise. If scores stall, the system performs a reflection step to categorize remaining failures, identifying issues like flawed graders or outdated API priors.
Benchmark Results
In an internal customer support benchmark, the workflow reduced costs significantly while improving accuracy. Starting from a baseline of Opus 4.8 at High effort (74.4% accuracy, 4.6 cents/ticket), the process optimized to Sonnet 5 at Low effort, achieving 98.9% accuracy at approximately 1 cent/ticket. On a held-out test set, the final configuration reached 90.5% accuracy compared to the original 78.6%, at roughly one-fifth the cost.
For the claude-api skill itself, hillclimbing improved accuracy from a 66% baseline to 87.9% over 24 rounds. Key improvements included adding coverage for missing features, fixing type table errors, and correcting graders that contradicted documentation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.