DelusionEval measures psychological harm and delusional spirals in LLMs
August 6, 2026
DelusionEval introduces an evaluation protocol using 12,591 messages from 18 participants to test if LLMs promote user delusions. Findings indicate that model size, release date, and test-time reasoning do not reliably correlate with the tendency to exhibit delusion-linked behaviors.
HOW THIS AFFECTS YOU
●
researcherYou can use this protocol to evaluate how context length influences a model's tendency to reinforce psychological harm.
●
policyThis provides a new framework for assessing the safety and mental health risks of chatbot interactions.