CUDA Implementation of B200 Attention Kernel Reaches 94.4% of FA4
September 2, 2026
A new implementation of a dense B200 attention kernel using CUDA and PTX achieves 94.4% of FlashAttention-4 performance across 4K, 8K, and 16K shapes. The project includes a visual guide for optimizing kernels on Blackwell hardware.
HOW THIS AFFECTS YOU
●
builderYou can use this guide to implement high-performance attention kernels for Blackwell architecture.
●
researcherThis provides a benchmark for efficient attention mechanisms on the latest NVIDIA hardware.