Issues / #1468
#1468 prefill copy_i32: illegal memory access on long prompts when the prefill chunk is large (auto:32768)
open · @btc1000w · 1 commentaires · Sur GitHub
Server & APIMulti-GPUNVIDIA / CUDAModels & quantsWindows
Description
## Summary
`prefill copy_i32` faults with an **illegal memory access** on **long prompts when the prefill chunk is large**, then `check()` calls `std::exit(1)` and the engine dies (the server restarts it, ~100 s). It is **intermittent** (the same prompt sometimes succeeds) and **reproduces on both 0.1.40 and 0.1.40.2**.
The chunk size is the isolated variable:
| `--prefill` | actual chunk | requests | crashes |
|---|---:|---:|---:|
| `auto:32768` | 19,200-20,480 | 45 | **11 (~35 %)** |
| `8192` | 8,192 | 20 | **0** |
| `auto:16384` | 16,384 | 16 | **0** |
Removing `--pipeline-windows` did **not** stop it (1 crash in 4 requests), so it is not that path.
## Where it fails
```
src/prefill/kernels.cu:70 void check(const char* what) {
const cudaError_t e = cudaGetLastError();
if (e != cudaSuccess) { std::fprintf(stderr, "prefill %s: %s\n", what,
cudaGetErrorString(e)); std::exit(1); }
}
src/prefill/kernels.cu:1677 __global__ void copy_i32_kernel(int32_t* dst, const int32_t* src, int64_t n)
src/prefill/kernels.cu:1695 void copy_i32(int32_t* dst, const int32_t* src, int64_t n, void* stream) {
... copy_i32_kernel<<<min((n+255)/256, 256), 256, 0, stream>>>(dst, src, n);
check("copy_i32");
}
```
Call sites (all pass `T * K` int32, `T` = chunk, `K` = top-k):
```
src/prefill/prefill.cpp:2694 if (grp_mapped) copy_i32(m.grp_dev, m.ids, T * K, m.cs);
src/prefill/prefill.cpp:2829 copy_i32(m.slot_dev, m.grp_dev + m.grp_tk, T * K, m.cs);
src/prefill/prefill.cpp:2830 copy_i32(m.src_dev, m.grp_dev + 2 * m.grp_tk, T * K, m.cs);
```
The engine log line is exactly:
```
prefill copy_i32: an illegal memory access was encountered
```
then the server logs `engine started: ...` (a fresh engine), so the request is lost and the model reloads.
## Reproduce
Engine 0.1.40 (and 0.1.40.2), Windows:
```
strata.exe --serve --pack <iq3_xxs pack> --native <shard1> --ple-gguf <shard2> \
--expert-profile data/expert-profile.bin --expert-cache auto \
--prefill auto:32768 --spec 4 --spec-min-p 0.70 --mtp <mtp/rt> \
--max-context 262144 --kv int8 --kv-resident 32768 --vision \
--vram-reserve-mib 690 --pool-workers 8 --trim-stage-weights \
--pipeline-windows 2 --layer-split 26
```
Send a prompt of ~54 K-114 K tokens. The engine logs `prompt chunk auto: 19968 tokens, a 512-slot ring` at startup, then crashes within a few long-prompt requests. `--prefill auto:16384` (or `8192`) is stable.
## Environment
- Engine: 0.1.40 and 0.1.40.2 (`BUILD.json`, CUDA 13.0, archs 75/86/89/120 + PTX), `source: release`
- GPUs: RTX 3080 Laptop 16 GB + RTX 3070 Laptop 16 GB, driver 616.56, `--layer-split 26` (2 stages)
- CPU: Intel Xeon E5-2680 v4 (14c/28t), 95.8 GiB RAM, Windows
- Model: `Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS` (47.3 GB shard1 + 28.8 GB shard2, mixed expert types: IQ3_S/IQ4_NL/Q2_0/IQ2_S/IQ2_XS/IQ2_XXS/IQ3_XXS)
- Pack: `native_experts.txt` regenerated with `tools/iq_pack.py --skip-experts` for this image (so the expert offsets match); `dense.bin`/`index.txt` are byte-identical to before
- Other settings: `--trim-stage-weights`, `--pipeline-windows 2`, `--kv int8 --kv-resident 32768`, `--spec 4 --spec-min-p 0.70`, `STRATA_PREFILL_RING=1024`, vision encoder on the CPU (`vision.gpu=false`)
## Question
Is there a buffer in the prefill MoE path sized for a **smaller `T`** than the configured chunk, or a bound that `T * K` can exceed when the chunk is large? The failure being intermittent at the same prompt suggests a size/offset edge rather than a constant out-of-range pointer.
I can attach the engine's own stall dumps (`strata-stall-*.dmp`, written by the issue #29 watchdog) if that helps.Sur le site
Liens install, modèles, releases.