chore: relocate kernel benchmarks to csrc/bench
- Move benchmark_gemv.py, benchmark_swiglu.py, and benchmark_gemv_common.py from scripts/tools/ to csrc/bench/ so kernel benchmarks live next to the kernels they measure - Update reproduction commands in decode_linear_benchmark.md, swiglu_benchmark.md, and cuda_kernels.md - Codify the placement convention in AGENTS.md: kernel benchmarks in csrc/bench/, pure-CUDA harnesses in csrc/tests/*.cu, engine and evaluation benchmarks in scripts/
This commit is contained in:
@@ -84,7 +84,7 @@ OPT 1.3B M=1 is +4.54%. Qwen2 and LLaMA 3 70B M=1, and all three new
|
|||||||
families at M=8, remain exact PyTorch fallbacks.
|
families at M=8, remain exact PyTorch fallbacks.
|
||||||
|
|
||||||
These are synthetic projection-chain measurements, not whole-model throughput
|
These are synthetic projection-chain measurements, not whole-model throughput
|
||||||
claims. Reproduce them with `scripts/tools/benchmark_gemv_common.py`.
|
claims. Reproduce them with `csrc/bench/benchmark_gemv_common.py`.
|
||||||
|
|
||||||
The AstrAI 1B A→B→B→A results use the real `InferenceEngine`, including scheduler,
|
The AstrAI 1B A→B→B→A results use the real `InferenceEngine`, including scheduler,
|
||||||
sampling, and CUDA Graph. M=8 stays on PyTorch because its remaining
|
sampling, and CUDA Graph. M=8 stays on PyTorch because its remaining
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Decode linear shape benchmark
|
# Decode linear shape benchmark
|
||||||
|
|
||||||
`scripts/tools/benchmark_gemv.py` records the `F.linear` baseline used to decide
|
`csrc/bench/benchmark_gemv.py` records the `F.linear` baseline used to decide
|
||||||
whether a BF16 GEMV or small-M kernel should enter automatic inference dispatch.
|
whether a BF16 GEMV or small-M kernel should enter automatic inference dispatch.
|
||||||
It does not change model execution or select a custom kernel.
|
It does not change model execution or select a custom kernel.
|
||||||
|
|
||||||
@@ -10,7 +10,7 @@ modes. Results include device-event latency samples, p50/p90/p99, estimated
|
|||||||
effective IO bandwidth, and CUDA kernel launches per call.
|
effective IO bandwidth, and CUDA kernel launches per call.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CUDA_VISIBLE_DEVICES=0 python scripts/tools/benchmark_gemv.py \
|
CUDA_VISIBLE_DEVICES=0 python csrc/bench/benchmark_gemv.py \
|
||||||
--output results/decode_linear.json \
|
--output results/decode_linear.json \
|
||||||
--markdown-output results/decode_linear.md
|
--markdown-output results/decode_linear.md
|
||||||
```
|
```
|
||||||
@@ -25,7 +25,7 @@ For direct A/B coverage of the custom kernel and guarded dispatcher across
|
|||||||
traditional LLaMA and GPT-NeoX decode shapes, use:
|
traditional LLaMA and GPT-NeoX decode shapes, use:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CUDA_VISIBLE_DEVICES=0 PYTHONPATH=. python scripts/tools/benchmark_gemv_common.py \
|
CUDA_VISIBLE_DEVICES=0 PYTHONPATH=. python csrc/bench/benchmark_gemv_common.py \
|
||||||
--suite all --family traditional --m 2 4 \
|
--suite all --family traditional --m 2 4 \
|
||||||
--output results/gemv_common.json
|
--output results/gemv_common.json
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -1,12 +1,12 @@
|
|||||||
# Fused SwiGLU benchmark
|
# Fused SwiGLU benchmark
|
||||||
|
|
||||||
`scripts/tools/benchmark_swiglu.py` compares the directly callable fused BF16
|
`csrc/bench/benchmark_swiglu.py` compares the directly callable fused BF16
|
||||||
SwiGLU primitive with both `F.linear` and the existing two-GEMV chain. It covers
|
SwiGLU primitive with both `F.linear` and the existing two-GEMV chain. It covers
|
||||||
the native AstrAI 1B MLP plus LLaMA 2 7B/13B, LLaMA 3 8B, and GPT-NeoX 20B
|
the native AstrAI 1B MLP plus LLaMA 2 7B/13B, LLaMA 3 8B, and GPT-NeoX 20B
|
||||||
up/gate shapes at M=1/2/4/8 in eager and CUDA Graph modes.
|
up/gate shapes at M=1/2/4/8 in eager and CUDA Graph modes.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
CUDA_VISIBLE_DEVICES=0 python scripts/tools/benchmark_swiglu.py \
|
CUDA_VISIBLE_DEVICES=0 python csrc/bench/benchmark_swiglu.py \
|
||||||
--output results/swiglu.json \
|
--output results/swiglu.json \
|
||||||
--markdown-output results/swiglu.md \
|
--markdown-output results/swiglu.md \
|
||||||
--m-values 1,2,4,8 --mode both \
|
--m-values 1,2,4,8 --mode both \
|
||||||
|
|||||||
Reference in New Issue
Block a user