SubSegGPT and SubSegDeBERTa Learn Tokenization During Pretraining
September 2, 2026
Subword segmental language modeling moves tokenization from a preprocessing step to a learned component during training. SubSegDeBERTa and SubSegGPT demonstrate improved sample efficiency and zero-shot performance by discovering subword units optimized for the specific training objective.
HOW THIS AFFECTS YOU
●
researcherYou can explore learned tokenization to improve sample efficiency in small-scale language models.