[arXiv]score: 0.18
Through the Looking Glass: Directly Reading and Writing Transformers
September 10, 2026
Analyzing 18 models from 124M to 7B parameters reveals that token predictions rely on a tiny fraction of total components, often fewer than 16. While thousands of units contribute to logits, net contributions are driven by a small sufficient set, with 75% of layer updates acting as fixed linear maps.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy