ViperEkura
15862d4b56
perf: fuse fp8 linear fwd and bwd into single kernel calls
- fp8_linear_forward: cast + cublasLt GEMM + transpose + bias in one call
- fp8_linear_backward: scale-free, dtype derived from input tensor
- drops per-op Python dispatch (was ~6-8 launches per linear) and amax syncs
- 1024x1024 linear: 6.8x slow -> 0.67x (36.7us vs 24.8us bf16)
- small-model e2e still 1.71x slow; 15bt estimate ~0.78x (linear-heavy)
2026-08-14 01:24:37 +08:00
..
2026-08-09 11:38:50 +08:00
2026-08-08 16:15:10 +08:00
2026-08-14 01:24:37 +08:00
2026-08-09 13:32:40 +08:00
2026-08-08 23:43:05 +08:00
2026-08-01 09:20:54 +08:00
2026-07-31 08:32:22 +08:00
2026-08-05 22:20:29 +08:00
2026-07-29 12:50:27 +08:00
2026-07-30 09:38:20 +08:00
2026-08-09 13:40:27 +08:00
2026-08-08 12:39:27 +08:00
2026-07-30 09:38:20 +08:00
2026-08-08 23:43:05 +08:00
2026-05-24 20:35:44 +08:00
2026-07-28 10:36:17 +08:00