SparsePR Method for Training-Free Sparse Attention
August 18, 2026
SparsePR introduces Response-Coupled Partitioning and Probe-Fitted Residual Reconstruction to accelerate video transformers without retraining. It uses sampled-query key responses to form paired K/V groups and an affine correction to minimize post-softmax error from skipped interactions.
HOW THIS AFFECTS YOU
●
researcherYou can potentially reduce inference latency in video generation models without the cost of retraining.