NameTrace reveals unequal token access for diverse names
September 27, 2026
The NameTrace framework shows that LLM tokenizers provide unequal lexical access to names based on race and gender metadata. Many names are fragmented into multiple subwords, while others receive single-token access, potentially influencing model behavior.
HOW THIS AFFECTS YOU
●
researcherThis provides a framework to measure pre-behavioral biases at the lexical interface.
●
policyYou should investigate how tokenizer bias contributes to systemic algorithmic unfairness.