●researcherYou can use this benchmark to evaluate whether models truly reason or just recall drug associations.
●policyThis provides evidence for the necessity of rigorous, counterfactual testing in medical AI regulation.
●healthThis highlights critical safety risks in using LLMs for clinical decision support.