- Removes per-step STAGES*BC*LD smem clear loop (2 buffers x 24 layers) - cp.async predicated load + softmax mask already exclude padding slots, matching the paged prefill kernel which never zero-inits - Standalone and extension tests pass; decode step time unchanged