AivexaNewsSearch
AI news for builders and product teamsChecked every hour
PyTorchFirst partyResearch

Low Precision Flash Attention 4: End-to-End Block-Scaled Attention for Blackwell

Collected Oct 1, 2026

PyTorch detailed an extension of FlashAttention-4 (FA4) with end-to-end MXFP8 support for both forward and backward passes on Blackwell hardware. The work, attributed to Meta, reports 2.85 PF/s forward and 2 PF/s backward on LLM shapes, and on internal shapes 2.54 PF/s forward and 1.58 PF/s backward, described as up to 1.6x and 1.52x gains over BF16.

The implementation targets Blackwell's block-scaled MMA instructions (tcgen05.mma.block_scale), which operate natively on microscaling formats including MXFP8, MXFP6, MXFP4 and NVFP4. The report describes four contributions: TMEM allocation strategies that fit scale factors into the fully utilized 512-column TMEM with minimal new barriers; online, transpose-invariant square block scale quantization of dS using Blackwell's redux.sync.max.abs.f32 warp-wide reduction; fused RMSNorm+Quantize and GEMM+Quantize kernels producing FP8 output with dual scale factor layouts in one pass; and a zero-gather jagged module where FP8 data stays at unpadded positions while only the smaller scale factors are scattered, padded and permuted to 128-aligned addresses for TMA.

For the forward pass, the softmax warp computes softmax in FP32 and converts results to MXFP8 while computing scales, reusing the max already computed for softmax. The backward pass quantizes Q, K and dO in square [32,32] blocks to make the E4M3 payload transpose-invariant, so both GEMMs can reuse one quantized representation. The report states dQ reduction was identified as a throughput bottleneck and that FP16 dQ with static scaling yields optimal performance, with dQ values described as consistently small in production workloads.

According to the report, the module is used internally at Meta for GEM training, and the authors describe it as one of the first SoTA implementations of MXFP8 FA4 forward and backward used in production training workloads. The code has been open sourced in the facebookresearch/ads_model_kernel_library repository under the lp_fa4 path. The report also notes a 1.76 PFLOP/s jagged GEMM-Norm backward result that is treated as performance-only, not an accuracy-qualified comparison, and states plans to open a CuDNN GitHub issue to investigate the matter.

Read at PyTorch

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

TL;DR We extend FlashAttention-4 [1] with MXFP8 forward and backward, reaching 2.85 PF/s forward and 2 PF/s backward on LLM shapes. On our internal shapes, FA4 MX8 reaches 2.54 PF/s...