[arXiv]score: 0.18
More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
October 6, 2026
Sparse Asymmetric Group-Query Attention (SAGA) decouples key and value head counts to reduce query-key computational bottlenecks during LLM decoding. Combined with approximate top-N attention, this method allows for reduced key heads while maintaining value head capacity to preserve model expressivity without increasing latency.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy