Issues / #1756
#1756 [Bug]: CPU prefill share: a failed cudaHostAlloc kills the engine instead of falling back to the GPU-only path (single GPU, Windows, 64 GB host)
open · @noderex · 0 comentarios · En GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
What happened
A long prompt killed the engine. It reported prefill: cannot allocate the CPU experts' activations, the process exited, and the Python server restarted it (the request itself failed with 0 tokens).
The failing allocation is the CPU prefill share's pinned activation buffer (cudaHostAlloc, src/prefill/prefill.cpp:3006). The share is ON by default in 0.1.41, and on this host every pinnable page was already taken by the arena, so the request died in an optional optimization — the GPU-only path (STRATA_PREFILL_CPU_SHARE=0, "identical to 0.1.40.3" per the source comment) would have run it.
Prompt that triggered it: 41,838 tokens — 5 × 8,192-token chunks read fine, then the 878-token tail (< 1024, so the share arms on it) needed a bigger pinned block and the allocation failed.
Engine log, the last lines before the process left:
代码块
strata serve: decode expert cache hit rate: 98.3% (144824 hits / 147303 lookups); 57 more read by the GPU over PCIe or from another GPU (0.0% of all 147360 routed)
strata serve: KV streaming: 99.51% of 21745740 block reads hit VRAM, 430.9 MiB read from RAM
strata serve: suffix drafts: 1 windows, 5 of 5 drafts accepted
strata serve: prefill: cannot allocate the CPU experts' activations
The server's own view (serve/server.py narration):
代码块
[strata] reading the prompt: 40,960 of 41,838 tokens, 6 s so far
[strata] the engine reported an error: prefill: cannot allocate the CPU experts' activations
[strata] done: 0 tokens in 6 s (0.0 tok/s) (error, cancel=False)
[strata] the engine stopped unexpectedly (exit code None). The usual cause is running out of RAM: ...
[strata] the engine had stopped (exit code None); starting it again (a minute or two) ...
[strata] the engine is running again
Recovery worked as designed (~70–100 s, no data loss). Everything else was healthy: the requests right before it ran 84.8–94.6 tok/s at 98.3–98.6% expert-cache hit.
Strata version, GPU, OS
0.1.41 (source build from fb58e0d, 2026-10-08) · RTX 3090 24 GB (sm_86) · Windows 10 19045 · single GPU · 64 GB RAM (63.94 GiB visible) · Xeon E5-2670, 8C/16T, AVX only (no AVX2/FMA/F16C/BMI2) · MSVC 14.44.35207 + CUDA 12.6, -DSTRATA_ISA_FLOOR=avx · model: Qwen3.8-Flash-Next IQ2_XS (self-packed) · Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS, 262,144-token context, --vision (CPU encoder).
Engine command line (from the log):
代码块
strata.exe --serve --pack <Strata-data>\packs\iq2_xs --native <...>-00001-of-00002.gguf
--ple-gguf <...>-00002-of-00002.gguf --expert-profile <...>\data\expert-profile.bin
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <Strata-data>\mtp\rt
--max-context 262144 --kv int8 --kv-resident 32768 --vision --vram-reserve-mib 1200
The GPU was not the limit: 12215 slots, 16.37 GiB of VRAM, 909 MiB of VRAM free with everything loaded, no paging indicators (ram_blobs = file_blobs = file_mb = 0 on every request).
Host state at the time (this is where it ran out)
代码块
physical free: 10.81 GiB of 63.94 GiB
system commit: 80.55 GiB of 107.94 GiB (63.94 RAM + 44 GiB page file)
strata.exe private: 61.21 GiB commit, 34.45 GiB working set
The startup log had already said there was no pinnable memory left:
代码块
strata generate: expert arena: locked 5901 MiB via working-set minimum + VirtualLock; cudaHostRegister of the
whole arena FAILED (out of memory); 40 slices pinned (27 GiB); large pages refused for
35456548864 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata generate: KV streaming: 32768 of 262144 cells per QSA layer in VRAM, the K/V in 3.09 GiB of pinned RAM
So the host is holding ~27 GiB of pinned arena slices + 3.09 GiB of pinned K/V + 5.9 GiB VirtualLock, on a 64 GB box, and the prefill then asks for one more pinned block.
Where it comes from
src/prefill/prefill.cpp:3006, inside the share branch (if (cpu_maybe && cpu_arm)), growing the buffer:
cpp
const size_t want = (size_t) T * N;
if (m.cpu_x_n < want) {
cpu_cold = true;
if (m.cpu_x) cudaFreeHost(m.cpu_x);
m.cpu_x = nullptr;
m.cpu_x_n = 0;
if (cudaHostAlloc((void**) &m.cpu_x, want * sizeof(float), cudaHostAllocDefault) != cudaSuccess) {
err = "prefill: cannot allocate the CPU experts' activations";
return false; // <- the whole prefill fails, the engine exits
}
m.cpu_x_n = want;
}
Both messages (…activations at :3006 and …rows at :3183, the second one also cudaHostAlloc) exist only in the share paths, so the error text identifies the code that failed without ambiguity.
The share is on by default in 0.1.41 (prefill.cpp:157), and it arms for chunks below STRATA_PREFILL_CPU_SHARE_MAX (default 1024, prefill.cpp:177-185) — so any short prompt and the last sub-chunk of a long prompt go through this allocation. On this host the preceding prompt in the same session ended with a 5-token tail; the one that died ended with an 878-token tail.
The buffer grows by free-then-allocate, so each new maximum chunk size releases and re-requests a pinned block.
The single-GPU path has no pin cap. generate.cpp:4217-4221 computes pin_wddm_cap only when expert_cache_remote[0] > 0 || multi_gpu, so on a one-GPU box the arena goes straight for the whole-arena cudaHostRegister; when the driver refuses it, the sliced fallback pins what it can (27 GiB here) and nothing keeps a reserve for later allocations. The comment right above that code describes this same family of failure ("pinning all of it … leaves WDDM refusing every later allocation — measured on the 5080 + 3090 rig: cudaMemGetInfo and the next cudaMalloc fail"), but the guard only covers multi-GPU.
The Linux path has pacing for exactly this class of problem (#1250: STRATA_PIN_GUARD, STRATA_PIN_RESERVE_GIB, expert_source.cpp:2374-2390), but that block is #if defined(__linux__).
Why I think this is a bug, not a setting
The share is an optimization that the engine can run without (its own comment: off = byte-identical to 0.1.40.3). Failing it should mean "share off for this run", not "the engine exits and the user's request dies".
The engine had already been told there was no pinnable memory at startup (cudaHostRegister … FAILED (out of memory)) and still armed a path that needs more of it — nothing carries that signal forward.
A host whose whole-arena registration failed is the strongest possible evidence that the pin budget is gone, yet the only automatic cap (8 GiB under WDDM) is gated on multi-GPU.
Suggested fix (any one of 1–2 stops the crash; 3 removes the churn)
Degrade instead of failing. If cudaHostAlloc for cpu_x / cpu_rows fails, set the share off for the rest of the run (the same state STRATA_PREFILL_CPU_SHARE=0 produces), print one line, and keep serving the request on the GPU-only path. m.cpu_x/cpu_arm already model "running without the share".
Keep a pinnable floor. Treat a failed whole-arena cudaHostRegister as "budget exhausted": don't arm the share (or cap the arena pin / pace the registration with a reserve, like the Linux STRATA_PIN_GUARD path), so later cudaHostAlloc calls still have pages.
No mid-run churn. Grow cpu_x once to the largest chunk the run can produce (the chunk size is known before the first chunk) and never free-then-reallocate during a run.
Workaround I run now
In the model config, which the Python server passes to the engine's environment:
json
"env": { "STRATA_PREFILL_CPU_SHARE": "0" }
Worth noting for the default: the share's gain was measured on an RTX 5090 + 9950X3D (prefill.cpp:157-160), while this host is an AVX-only Xeon E5-2670, which the engine itself flags as the slow CPU path ("EXPERIMENTAL older-CPU build (ggml-cpu for avx) … expect it to be slow"). If the share is a loss here as well, its default could be gated on the CPU the way other things are.
What I did not test
Linux, multi-GPU, any host other than this one, any quant other than IQ2_XS.
The exact block size the failed call asked for (T * N floats) — I did not instrument it.
Whether more host RAM, a smaller arena, or STRATA_ARENA_PIN_GIB=<N> avoids it. I can run either and report.
The share's A/B on this CPU — no same-build baseline yet; I'll post the numbers in this thread.
It happened once; I have not made it deterministic. The trigger conditions I can name are: single GPU + Windows + a host whose arena pin fell back to slices (27 GiB pinned of 64 GB) + a share-eligible chunk (< 1024 tokens) larger than any earlier one in the run. Say the word and I'll try to reproduce it with the share on and send the full strata-<model>.log.En el sitio
Enlaces a install, modelos, releases.