- 128x128 CTA of 8 warps x 64x32 warp tiles: 16 mma.sync per warp per K-segment (was 8) - ldmatrix.x4/x2 with per-lane swizzled addresses replaces 36 scalar LDS per warp-tile step - __launch_bounds__(256, 2) caps registers at 124 so two CTAs fit per SM - fp8 linear vs bf16 cuBLAS: fwd 1.07x->2.65x, bwd 1.41x->2.65x by size, peak 31-33 TFLOPS