Anthropic added two new commands to its claude-api skill for Claude Code that let developers build a test set for their AI application and then have Claude iterate against it on its own. The company described the tools, called build-eval and hillclimb, in a September 28 blog post credited to Lance Martin.

The pitch addresses a problem every team shipping an AI feature runs into eventually: knowing whether a change actually made the product better, or just better at the ten examples someone happened to test. Anthropic’s post lays out what it considers a well built evaluation. Tasks should mirror real production use rather than whatever is easy to generate or grade, and a strong model at high effort should still score well below 100 percent, since a test with no headroom cannot show whether a change helped.

Build-eval walks a developer through picking sample cases, prioritizing real production transcripts and bug reports over hand-written or synthetic ones, then proposes a grader: a simple pass-fail code check when outputs are constrained, or a second model acting as judge when the answer space is open-ended. Anthropic says the tool pauses for the developer to review both the cases and the grader’s scores before running anything at scale.

Hillclimb is the half that acts on the results. Given an evaluation, it splits cases into a training set and a held-out test set, proposes one change per round (a prompt edit, a different model, a new tool description), reruns the eval, and keeps the change only if both sets improve. If a change lifts the training score but leaves the test score flat, Anthropic says Claude treats that as a sign of overfitting and reverts it. The company frames this as guarding against a system that looks better on its own benchmark while doing nothing different in production.

Anthropic backs the pitch with two internal examples rather than independent benchmarks. In a customer-support test with 44 tickets, the company says switching from Opus 4.8 to Sonnet 5 on a low effort setting, combined with prompt fixes the tool made itself, took accuracy from 74.4 percent to 98.9 percent while cutting token cost to roughly a fifth of the starting price. In a second case, Anthropic says hillclimb raised its own claude-api skill’s evaluation score from 66 percent to about 88 percent by finding coverage gaps and outdated code examples the skill was still recommending. Both figures come from Anthropic’s own testing, and the release does not include a comparison against a similar tool from another lab or an independent replication.

The company is effectively productizing a workflow that engineering teams at OpenAI and Google DeepMind already run informally with human researchers driving the eval-and-tune loop by hand. What Anthropic is selling here is turning that loop into a repeatable command a developer runs from inside Claude Code, which matters more for how fast a team can iterate than for what any single number proves. Whether the overfitting check, splitting train from test and reverting changes that do not generalize, holds up on messier real-world data than Anthropic’s two examples is something teams will only learn by running it themselves.

Teams that already maintain an internal eval harness for an AI feature should try pointing hillclimb at it on a low-stakes prompt change first, and check whether its test-set score matches what a human reviewer would say about the same output.

Reported from Anthropic’s Claude developer blog (claude.dev), published September 28, 2026.