- 64-token prefill forward triggers cuBLAS auto-tuning at init - reduces first-chat prefill from ~520ms to ~27ms - warmup decode also drops from ~215ms to ~71ms
- 64-token prefill forward triggers cuBLAS auto-tuning at init - reduces first-chat prefill from ~520ms to ~27ms - warmup decode also drops from ~215ms to ~71ms