[arXiv]score: 0.24
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
August 18, 2026
A psychophysics-inspired benchmark evaluates LLM diagnostic accuracy and confidence calibration using synthetic Alzheimer-type versus depression-related neurocognitive vignettes. Pilot results with gpt-4.1-nano achieved 93.5% accuracy and an AUROC of 0.876, showing that model confidence tracks evidence strength and information quality.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy