●researcherYou can use these benchmarks to evaluate the calibration and safety of models in specialized scientific domains.
●policyThis highlights the risks of model hallucinations and the importance of reliable abstention in high-stakes scientific applications.