RBS-Attention is a training-free sparse-prefill method that uses a centroid base branch and a maximum key-block radius rescue branch to prevent mean dilution. On H100 GPUs, it achieves up to 20.65x standalone prefill-attention speedup.
HOW THIS AFFECTS YOU
●
builderYou can significantly reduce latency for long-context applications without retraining models.
●
researcherThis provides a mechanism to handle highly relevant tokens in sparse attention without losing precision.