Fixed Token Codes Achieve Competitive Performance Without Trainable Embeddings
October 6, 2026
Research shows 1.7B-class decoder-only models can achieve substantial capability using fixed token-ID codes or GF(2) recoding instead of learned embedding tables. Models using canonical 16-bit codes reached 52.40% HellaSwag and 70.51% PIQA accuracy during 100B token training.
HOW THIS AFFECTS YOU
●
builderYou could potentially reduce model parameter counts and memory overhead by using fixed input interfaces.
●
researcherThis suggests token-specific parameterization may not be strictly necessary for core language modeling capabilities.