refactor: remove bf16 gemm and swiglu kernels and rebuild csrc benchmarks

- delete csrc/kernels/gemm.cu and swiglu.cu and drop their CMake and setup.py registration
- remove the ops wrappers plus backend/linear.py and backend/swiglu.py so Linear and MLP call F.linear directly
- drop the four gemm and swiglu kernel test files and prune the stale cuda_kernels.md sections
- add csrc/bench benchmarks for the remaining kernels: attention decode prefill paged decode paged prefill versus single-launch SDPA references, rotary versus the torch fallback, fp8 quantize and mm_fp8 versus torch baselines
- attention, rotary_emb, and fp8_ops kernels are unchanged
This commit is contained in:
2026-09-05 01:38:10 +08:00
parent a77e35dd51
commit 6709534d64
24 changed files with 1306 additions and 3394 deletions
-4
View File
@@ -11,9 +11,7 @@ from astrai.extension.backend.attention import (
attn_backend,
get_backend,
)
from astrai.extension.backend.linear import linear
from astrai.extension.backend.rotary import apply_rotary_emb
from astrai.extension.backend.swiglu import swiglu
__all__ = [
"ATTN_BACKEND",
@@ -26,6 +24,4 @@ __all__ = [
"attention",
"attn_backend",
"get_backend",
"linear",
"swiglu",
]