TopK-Guided Training-Free Activation Sparsity for Efficient LLM Inference
October 2, 2026
TopK-Guided improves LLM inference efficiency by combining token-level sparsity adaptation with sensitivity-aware budget allocation across transformer blocks. Testing on Llama-2 and Llama-3 shows improved perplexity and accuracy compared to fixed-budget or threshold-based methods while maintaining sparsity targets.
HOW THIS AFFECTS YOU
●
builderYou can implement more efficient, budget-aware inference without retraining models by optimizing activation sparsity across blocks.