Minima-KV achieves 3.5x KV cache compression with mixed-format paged attention
August 26, 2026
Minima-KV uses a hierarchical mixed-format approach, keeping anchor pages in FP8 and older pages in packed TQ3. This achieves 18.3 KiB per token on Qwen3.6-27B, providing 3.50x compression over BF16 while maintaining RULER needle-in-a-haystack performance.
HOW THIS AFFECTS YOU
●
builderYou can significantly increase serving capacity and reduce memory bottlenecks for long-context models.
●
researcherThe use of globally normalized online-softmax merge allows for heterogeneous decoding without dense shadow caches.