[arXiv]score: 0.18
Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
August 12, 2026
Quantile-Guided Density Estimation (QGDE) estimates the composition of hidden pretraining corpora by mapping token ID-to-ratio distributions from known datasets to target tokenizers. Testing on the SmolLM tokenizer yielded mean relative errors below 3.08% for category-level mixtures, enabling more accurate inference of training data distributions from released vocabularies.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy