LORE-KV: Training-Free KV Cache Eviction via Monte Carlo Estimation
October 7, 2026
LORE-KV estimates the future utility of prompt tokens by sampling short autoregressive continuations from a frozen model. It uses these sampled trajectories to score token importance through leave-one-out attention-output deletion cost, optimizing KV cache eviction without retraining.
HOW THIS AFFECTS YOU
●
builderYou can reduce memory overhead during inference without needing to train new eviction models.
●
researcherThis provides a training-free method for future-aware KV cache management.