HLA: Hybrid Linear Attention via Query-Dependent Chunk Mixing
October 4, 2026
Hybrid Linear Attention (HLA) improves long-context decoding by using content-dependent routing gates to interpolate historical affine state transitions with an identity map. This allows the model to adaptively access sparse or distant information within Gated DeltaNet (GDN) architectures.
HOW THIS AFFECTS YOU
●
builderYou can implement more efficient autoregressive decoding for long-sequence tasks.
●
researcherYou can achieve better selective access in long-context models using query-dependent chunk-level attention.