TypedBench Benchmark for Non-Generative Decision Models
October 9, 2026
TypedBench evaluates System One models on calibration, framing sensitivity, and cost-effectiveness across categorical, ordinal, and binary outcomes. It measures accuracy across paraphrases and calibration error relative to a matched, perfectly calibrated predictor.
HOW THIS AFFECTS YOU
●
builderYou can use this to ensure your probabilistic models don't trigger unintended software actions due to wording sensitivity.
●
researcherThis provides a more rigorous way to assess calibration error against the finite-sample noise floor.