- dX/dW run as fp8 cublasLt gemms via fused transpose-cast - shared (m,k,n) algo cache for fwd/bwd, mutex-protected - bias add in-place on bf16 output, drop output copy
- dX/dW run as fp8 cublasLt gemms via fused transpose-cast - shared (m,k,n) algo cache for fwd/bwd, mutex-protected - bias add in-place on bf16 output, drop output copy