●researcherUse this benchmark to ensure your clinical models aren't just learning to predict outcomes from historical labels.
●policyThis provides a framework for evaluating the safety and reliability of LLMs in high-stakes medical decision support.
●healthThis reveals why models may appear more competent in retrospective testing than in live clinical settings.