Autonomy-of-Heads Enables Data-Free Sparse Attention via Spectral Geometry
August 10, 2026
Autonomy-of-Heads identifies retrieval and streaming heads by analyzing the effective rank of query-key projection spectral geometry. This data-free method allows for sparse attention and KV-cache compression without relying on costly runtime attention scores or calibration prompts.
HOW THIS AFFECTS YOU
●
builderYou can reduce KV-cache costs and inference latency in long-context models without needing calibration data.
●
researcherThis offers a new weight-space method for diagnosing head functionality via the kernel attention operator.