TRACE Benchmark for Safety Evaluation of Reasoning Traces
August 26, 2026
TRACE introduces an evidence-grounded benchmark to evaluate safety across the entire inference pipeline, including prompts, reasoning traces, and final responses. It addresses the limitation of existing benchmarks that ignore unsafe content hidden within intermediate Large Reasoning Model (LRM) traces.
HOW THIS AFFECTS YOU
●
researcherUse this to evaluate if your reasoning models are generating unsafe internal logic despite safe outputs.
●
policyThis enables more rigorous safety audits of autonomous reasoning processes.