●builderYou must implement more specific, fact-checking-focused rubrics to accurately assess medical model safety.
●policyYou should be wary of relying on standard LLM benchmarks for safety and clinical compliance.
●healthYou need more granular, clinician-driven error injection to validate medical AI reliability.