MILO Enables Efficient Many-Shot ICL via Block-wise KV Compression
September 25, 2026
MILO uses block-wise low-rank compression to mitigate the linear KV cache scaling bottleneck in many-shot in-context learning. It dynamically allocates rank budgets based on information entropy to preserve the fidelity of high-importance context blocks.
HOW THIS AFFECTS YOU
●
builderYou can support thousands of demonstration examples in long-context windows with lower memory overhead.
●
researcherYou can exploit low-rank redundancy in KV caches to optimize many-shot inference efficiency.