Transformer Architectures Exhibit Linear Superposition of Input Streams
September 23, 2026
Transformers exhibit a Superposition Linearity Hypothesis where combining distinct text streams results in a superposition of next-token distributions. This intrinsic architectural property diminishes during pretraining but can be restored via lightweight fine-tuning to minimize distribution divergence.
HOW THIS AFFECTS YOU
●
researcherYou can leverage fine-tuning to recover linear properties for multi-stream processing tasks.