WhiteMatter allows Transformer layers to draw on past-token representations from any depth using a learned mixer for shared key-value (KV) cache channels. This architecture achieves performance comparable to standard Transformers with 50% more layers while using half the KV cache.
HOW THIS AFFECTS YOU
●
builderYou can deploy more capable models with significantly lower memory footprints and KV cache requirements.
●
researcherThis introduces a new method for optimizing information reuse across different model depths.