Issues / #1650

#1650 prefill: the batched prompt path aborts (cublasGemmEx f16, cuBLAS status 14) on chunks below about 32 tokens

open · @W1nge · 1 commentaires · Sur GitHub

Server & APINVIDIA / CUDAModels & quantsWindows

Description

With very small `--short-read` values the prompt path cannot serve anything: the first request aborts the engine.

```
prefill gemm: cublasGemmEx f16: cuBLAS status 14
```

(the engine exits with code 1, the server restarts it, the same thing happens to the next request).

`--short-read 0` sends every prompt segment to the batched path, including the last segment of a chat prompt – the role header the server appends, which is about 6 tokens. Chunks that small abort. Measured on one Windows host (CUDA 12.4.131, RTX 2080 Ti sm_75 + Tesla P100 sm_60, driver 537.13, native IQ3_XXS pack, 8192 context) with `STRATA_TRACE=1`, which prints the chunk before it runs:

| prompt chunk | result |
| --- | --- |
| 6 tokens (the turn header, `prompt chunk 0 of 6`) | aborts |
| 21 tokens | aborts |
| 24 tokens | aborts |
| 33 tokens | reads normally (2.30 s) |
| 37 tokens | reads normally |
| 73 / 109 / 195 / 429 / 546 / 877 tokens | read normally |

So the boundary sits between 24 and 33 tokens of chunk. It is not the first request that fails: after a chunk that reads, the next small chunk still aborts, and a session that reads anything through the verify windows first (any `--short-read` of 32 or more on these prompts) never shows it.

Minimal reproduction, no patches, stock 0.1.41 (`fb58e0d`) on Windows:

```
strata.exe --serve --short-read 0 ...      # any full-size serve configuration
```

then send one ordinary chat request whose prompt ends with the role header, e.g. a message of about 40 tokens. `STRATA_TRACE=1` makes the last lines

```
strata trace: request 40 0
strata trace: prompt chunk 0 of 33          <- reads normally
strata trace: prompt chunk 0 of 6           <- aborts here
prefill gemm: cublasGemmEx f16: cuBLAS status 14
```

The failing call is `Gemm::f16` in `src/prefill/gemm.cu` (`cublasGemmEx`, `CUDA_R_16F` operands, `CUBLAS_COMPUTE_32F`), which is the fp16 expert GEMM of the chunk (the row count is the number of experts the chunk routes to, `src/prefill/prefill.cpp` around the `dq_gu`/`dq_d` calls). Small chunks therefore reach it with very few rows.

I first hit this while measuring whether the batched prompt path can replace the verify windows for medium prompts: it makes `--short-read 0` unusable, and it means the minimum is implicitly about 32 tokens rather than the documented "0 = never use the windows". A guard (clamp the flag, or keep the windows for a segment the batched path cannot serve) or a fix in the small-chunk GEMM path would both be fine; the default `--short-read 64` is unaffected.

Sur le site

Liens install, modèles, releases.