[arXiv]score: 0.13
Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
September 11, 2026
Using OLMo-2 and Pythia datasets, this study quantifies membership evidence by replacing estimated training data with exact duplication counts. Evaluation of models from 1B to 13B parameters shows a rank correlation of -0.08 between model prediction costs and training exposure, suggesting low memorization signals in ordinary text.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy