Chunked KV-cache compression introduces phase sensitivity, where a token's position relative to compression windows causes retrieval performance to vary. In large open-weight models, this can lead to up to 40 percentage point differences in long-context retrieval accuracy across different phases.
HOW THIS AFFECTS YOU
●
builderBeware of inconsistent retrieval performance when implementing windowed KV-cache optimizations in production.
●
researcherYou should account for phase-based performance gaps when evaluating long-context compression methods.