Frontier Models Fail to Accurately Judge Psychological Safety
August 27, 2026
A study using aipsy-bench shows that frontier models like GPT-5.4-mini and Claude-Sonnet-4-6 are unreliable judges for conversational AI psychological safety. Results indicate structured disagreement on critical metrics, with Gemini showing high leniency and failing to flag self-harm risks.
HOW THIS AFFECTS YOU
●
researcherThis highlights a critical need for specialized, psychologist-corrected evaluation benchmarks for agentic safety.
●
policyYou cannot rely on standard LLM-as-a-judge frameworks for high-stakes safety or mental health compliance.