The FTC framework uses sequential structured Tucker compression to approximate attention heads while respecting Q/K/V structures. It achieves the lowest WikiText-2 perplexity across seven decoder-only models from 6B to 32B without requiring fine-tuning or gradients.
HOW THIS AFFECTS YOU
●
builderYou can compress attention layers more aggressively without the cost of gradient-based recovery.
●
researcherThis approach accounts for representation shifts during post-training compression stages.