RaReCache Framework for Cross-Model KV Cache Reuse
October 9, 2026
RaReCache enables large target models to reuse KV caches from smaller source models via selective recomputation. The method uses rank disagreement to identify information-dense tokens that fail linear mapping, allowing for accurate decoding without full context prefilling.
HOW THIS AFFECTS YOU
●
builderYou can reduce latency in multi-model cascades or model-switching sessions by reusing existing caches.
●
researcherThe rank disagreement metric provides a way to quantify transfer failures in cross-model cache mapping.