Issues / #1650
#1650 prefill: the batched prompt path aborts (cublasGemmEx f16, cuBLAS status 14) on chunks below about 32 tokens
open · @W1nge · 1 commentaires · Sur GitHub
Server & APINVIDIA / CUDAModels & quantsWindows
Description
With very small `--short-read` values the prompt path cannot serve anything: the first request aborts the engine. ``` prefill gemm: cublasGemmEx f16: cuBLAS status 14 ``` (the engine exits with code 1, the server restarts it, the same thing happens to the next request). `--short-read 0` sends every prompt segment to the batched path, including the last segment of a chat prompt – the role header the server appends, which is about 6 tokens. Chunks that small abort. Measured on one Windows host (CUDA 12.4.131, RTX 2080 Ti sm_75 + Tesla P100 sm_60, driver 537.13, native IQ3_XXS pack, 8192 context) with `STRATA_TRACE=1`, which prints the chunk before it runs: | prompt chunk | result | | --- | --- | | 6 tokens (the turn header, `prompt chunk 0 of 6`) | aborts | | 21 tokens | aborts | | 24 tokens | aborts | | 33 tokens | reads normally (2.30 s) | | 37 tokens | reads normally | | 73 / 109 / 195 / 429 / 546 / 877 tokens | read normally | So the boundary sits between 24 and 33 tokens of chunk. It is not the first request that fails: after a chunk that reads, the next small chunk still aborts, and a session that reads anything through the verify windows first (any `--short-read` of 32 or more on these prompts) never shows it. Minimal reproduction, no patches, stock 0.1.41 (`fb58e0d`) on Windows: ``` strata.exe --serve --short-read 0 ... # any full-size serve configuration ``` then send one ordinary chat request whose prompt ends with the role header, e.g. a message of about 40 tokens. `STRATA_TRACE=1` makes the last lines ``` strata trace: request 40 0 strata trace: prompt chunk 0 of 33 <- reads normally strata trace: prompt chunk 0 of 6 <- aborts here prefill gemm: cublasGemmEx f16: cuBLAS status 14 ``` The failing call is `Gemm::f16` in `src/prefill/gemm.cu` (`cublasGemmEx`, `CUDA_R_16F` operands, `CUBLAS_COMPUTE_32F`), which is the fp16 expert GEMM of the chunk (the row count is the number of experts the chunk routes to, `src/prefill/prefill.cpp` around the `dq_gu`/`dq_d` calls). Small chunks therefore reach it with very few rows. I first hit this while measuring whether the batched prompt path can replace the verify windows for medium prompts: it makes `--short-read 0` unusable, and it means the minimum is implicitly about 32 tokens rather than the documented "0 = never use the windows". A guard (clamp the flag, or keep the windows for a segment the batched path cannot serve) or a fix in the small-chunk GEMM path would both be fine; the default `--short-read 64` is unaffected.
Sur le site
Liens install, modèles, releases.