New Claude API commands /build-eval and /hillclimb automate the generation of test sets and graders. These tools facilitate iterative prompt and model tuning while implementing safeguards to prevent overfitting to evaluation sets.
HOW THIS AFFECTS YOU
●
builderYou can automate the cycle of prompt engineering and evaluation to improve system reliability.
●
researcherUse these automated graders to accelerate the optimization of model settings and prompt structures.