Growing datasets dissolve link between training data and model output
August 18, 2026
Research shows that as training datasets scale, surgically removing specific examples becomes increasingly difficult because the relationship between training data and generated output weakens.
HOW THIS AFFECTS YOU
●
researcherThis suggests current data attribution and removal methods may fail at scale.
●
policyThis complicates regulatory efforts to enforce 'right to be forgotten' in large-scale models.