Foil Bridges Looped Transformers and Mixture-of-Experts
September 27, 2026
The Foil method enables Looped MoE by flattening experts to double the pool per layer and untying attention to provide unique parameters per pass. This allows fixed-size models to utilize parameters more fully through increased expert usage.
HOW THIS AFFECTS YOU
●
builderYou can push the capabilities of small-parameter models by increasing their effective computation per token.
●
researcherThis introduces a new architectural way to combine recurrence with sparsity.