Anthropic Adds Claude Code Workflows to Build AI Tests and Tune Apps Against Them
Developers can create evaluations and revise their apps one step at a time. Anthropic’s reported cost and accuracy gains come from a 44-ticket internal trial.
Loading page…
Developers can create evaluations and revise their apps one step at a time. Anthropic’s reported cost and accuracy gains come from a 44-ticket internal trial.
Listen to this story
Anthropic’s new Claude API workflows formalize an eval-first loop: developers can assemble reviewed test cases and grading rules, then let Claude Code propose application changes against a search set while protecting held-out cases. In an internal 44-ticket support benchmark, Anthropic reports search-set accuracy rising from 74.4% with Opus 4.8 to 98.9% with Sonnet 5 after model and prompt changes; reported token cost fell from 4.6 cents to about 1 cent per ticket. Results are not live-traffic evidence, and test-c,
Use `/claude-api build-eval` to propose cases from production conversations, bug reports, support tickets, developer examples, or synthesized codebase examples; developers review*
`/claude-api hillclimb` uses 30 tickets for search and 14 as a held-out test in Anthropic’s benchmark, and can reverse patches that fail to improve held-out results.
The workflow advises against merging gains too small to distinguish from normal score variation.
Claude Code can now help build a test for an AI application, then revise the application to score better. Anthropic has added two guided workflows to its Claude API skill for that job. The appeal is faster iteration; the catch is that a better score only helps if the test reflects the work the app actually does.
The first command, /claude-api build-eval, creates an evaluation inside the developer’s codebase. Claude Code asks about the application, proposes examples and pauses for approval. It looks first to production conversations, after asking about data retention and sensitive information, then to bug reports, support tickets, cases the developer writes and examples synthesized from the codebase.
The developer reviews every proposed input before the test runs. Cases chosen only because today’s model fails them may capture that model’s particular weaknesses, rather than the hard work users need done. Production traffic has a blind spot too: Anthropic notes that users may mostly try tasks they already expect the app to handle.
Claude then proposes a way to grade the answers: a code check for tightly defined outputs or a separate model applying a rubric to open-ended ones. The developer reviews sample grades. The workflow also checks whether the grader changes its verdict on the same answer and whether errors or timeouts could distort the baseline.
The second command, /claude-api hillclimb, works against an evaluation the developer has built. The developer chooses the goal—such as better performance or lower cost without losing performance—and what Claude may edit. Options include prompts, instruction files, tool descriptions, model settings and application code.
Claude splits the cases into a search set it can inspect and a held-out test set. It proposes one patch per round, runs the evaluation and reverses a change if the held-out score stays flat or falls. If a measured gain is too small to distinguish from normal score variation, the workflow advises against merging the change.
Anthropic tried the workflow on an internal customer-support benchmark of 44 tickets, using 30 for the search and holding back 14. The comparison below shows its reported search-ticket results, not measurements of live customer traffic.
The trial began with Opus 4.8 at high effort. The first changes removed mandatory tool calls, a scratchpad step and contradictory prompt rules. Switching to Opus 5.5 at low effort raised reported search-ticket accuracy to 87.8% while cutting token cost to 1.9 cents per ticket.
The workflow then tried Sonnet 5 at low effort. Anthropic says a further prompt change brought that model to 98.9% accuracy at roughly the same cost. These steps changed both the instructions and the model; the result is not evidence that prompt edits alone delivered the savings.
Keeping cases out of the optimizer’s view helps catch a familiar failure: a fix tailored to examples in the test rather than to real tasks. Anthropic gives the example of adding a tool that helps on benchmark cases but rarely helps in production. The held-out check can flag that mismatch within this evaluation. A developer still has to decide whether the tickets, grades and cost goal represent the workload that matters.
Loading discussion...
Join the conversation
Tell us which human review step matters most to you.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.