The proposed SA-PPG metric improves upon the flawed G-AP metric by using stratified per-question probability sampling to detect benchmark contamination. This method prevents over- and under-suppression of memorized data, providing a more accurate measure of a model's genuine capability.
HOW THIS AFFECTS YOU
●
researcherYou can use this to more accurately evaluate whether your models are reasoning or just memorizing.