- cublasLt C layout and buffer switched to CUDA_R_16BF, halving output bandwidth - downstream ops (RMSNorm etc.) keep matching bf16 dtype, fused kernels stay - numeric error unchanged (0.19% vs fp32 ref on quantized inputs)
- cublasLt C layout and buffer switched to CUDA_R_16BF, halving output bandwidth - downstream ops (RMSNorm etc.) keep matching bf16 dtype, fused kernels stay - numeric error unchanged (0.19% vs fp32 ref on quantized inputs)