ViperEkura
3067a8e1a6
feat: unify attention backend with multi-dim mask support
- Add attention() functional entry delegating to active backend
- GQA/MLA forward calls attention() instead of inline cache/SDPA
- CUDA kernels support 2D/3D/4D mask via mask_h_stride field
- CudaBackend.fwd_decode builds 2D padding mask for mixed seq_lens
- KVCache.max_len precomputed in bind_tasks to avoid GPU sync
- batch==1 decode short-circuits mask=None
- Split tests into conftest, test_backend, test_backend_equivalence, test_kernel_mask
- 440 tests pass, L20 decode 1.44-1.60x speedup vs torch native
2026-07-30 20:38:34 +08:00
..
2026-07-30 08:41:14 +08:00
2026-07-29 12:50:27 +08:00
2026-07-30 20:38:34 +08:00
2026-07-30 20:38:34 +08:00
2026-07-30 20:38:34 +08:00
2026-07-30 07:54:54 +08:00
2026-07-29 12:50:27 +08:00
2026-07-29 12:50:27 +08:00
2026-07-30 09:38:20 +08:00
2026-07-30 09:38:20 +08:00
2026-07-29 12:50:27 +08:00
2026-07-30 09:38:20 +08:00
2026-05-24 20:35:44 +08:00
2026-07-28 10:36:17 +08:00