●researcherYou can use these power-calibrated audits to more accurately determine if your model has seen benchmark data during training.
●policyThis provides a mathematical basis for evaluating the integrity and validity of AI safety and capability benchmarks.