Issues / #1353
#1353 stager-transient-regression
open · @dag08 · 0 comentarios · En GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
Descripción
# Prefill: the "transient source" heuristic makes RAM-resident Q8 prefill 4-7x slower
First bad commit: `361537b` — *"prefill: UD-Q4_K_XL prompts 2.8x faster (16K: 57 -> 160 tok/s on an RTX 5070) - 32 stager threads and a 128-deep ring when the experts are read from the GGUF in place; MMQ for its Q4_K / Q5_K / Q5_1 experts"*.
The heuristic is still in `v0.1.40.1` (`src/prefill/prefill.cpp:861-862`, unchanged).
## Environment
| | |
|---|---|
| GPU | NVIDIA GeForce RTX 5090, 32607 MiB (driver 580.173.02); a second card (RTX 2080 SUPER, 8192 MiB) is present and idle |
| System RAM | 376 GiB |
| Host -> device | 13.8 GB/s (engine's own PCIe probe) |
| Model | Qwen3.8-Flash-Next, Q8_0 — pack `qwen38-q8_0` plus the native `Q8_0-0000x-of-00006.gguf` shards |
| Engine args | `--resident-budget-gib 240 --expert-cache 1944 --prefill auto --kv int8 --kv-resident 32768 --max-context 262144` |
| OS | Linux |
The experts are **pinned in system RAM** (`--resident-budget-gib 240`), not read from the SSD during prefill.
## Symptom
A 128k-token prompt (123,885 tokens) on a fully warm page cache:
| Engine | Prefill | 128k prompt |
|---|---|---|
| `v0.1.31` | 390-427 tok/s | 4.8-5.3 min |
| `v0.1.32` | 107.5 tok/s | 19.2 min |
| `v0.1.33` | 109.3 tok/s | 18.9 min |
| `v0.1.34` | 95.4 tok/s | 21.7 min |
| `v0.1.37` | 53.9 tok/s | 38.3 min |
| `v0.1.40.1` | 57.8 tok/s | 35.7 min |
Every run: same machine, same model, same prompt, same seed (707), page cache preloaded with `vmtouch -t` before each start, KV streaming and expert-cache hit rates identical (78.01% / 35-38%). Only the engine binary differs.
## Bisection
`git bisect` in a detached worktree, `v0.1.31` good / `v0.1.32` bad, one variable at a time (each candidate built with `-DSTRATA_ENABLE_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=120`, measured with the same warm-cache harness):
| Commit | Description | tok/s | verdict |
|---|---|---|---|
| `505e5e9` | PLE: the n-gram table in FP8 E4M3 | 433.7 | good |
| `ee1e18e` | serve: model aliases (#297) | 426.7 | good |
| `5ba8c5b` | docs only (`docs/DETAILS.md`, `docs/UNSLOTH_Q4.md`) | — | good |
| `d058fde` | setup only (`setup.py`, docs, tests) | — | good |
| **`361537b`** | **prefill: 32 stager threads and a 128-deep ring when the experts are read from the GGUF in place** | **107.6** | **bad** |
| `5080b92` | Merge branch 'f-132-amd' | 87.2 | bad |
| `d434c02` | Merge branch 'f-132-q4' (`v0.1.32`) | 107.1 | bad |
`361537b` is the first bad commit.
## Mechanism
The commit adds this to `prefill.cpp` (identical in `v0.1.40.1`, lines 861-862):
```cpp
bool files = false;
for (int64_t l = 0; src != nullptr && !files && l < g.n_layers; ++l)
for (int64_t e = 0; !files && e < g.n_expert; ++e) files = src->transient(l, e);
const int threads = stv ? std::clamp(std::atoi(stv), 1, 32) : files ? 32 : std::max(2, std::min(4, hw / 4));
if (files && std::getenv("STRATA_STAGER_RING") == nullptr) m.stager->kRing = 4 * threads;
```
The comment states the intent: *"most of a chunk's blobs are page faults on the SSD, so the copies need many reads in flight"*. That is right for the target case — Unsloth's UD-Q4_K_XL read in place beyond its RAM budget.
The check, however, only asks **whether** an expert source is `transient()`, never **where the bytes come from**. Our Q8 setup answers `transient() == true` while every blob is already pinned in RAM. It therefore gets the SSD profile — 32 copy threads and a 128-deep ring — and on this machine that is 4x slower than the 4-thread / 16-deep default the same commit keeps "for every other source".
A second, smaller effect points the same way: `ring_slots()` chose 384 slots in `v0.1.31` (`g_pinned_share >= 0.9 ? 384 : 96`) and picks 1024 in `v0.1.37+` (`fused_ring() ? 1024 : 384`). The 1024 ring is sized for the fused Q2_0 layouts, not for a Q8 pack.
## Verified workaround (no rebuild)
All three switches already exist in `v0.1.37` and `v0.1.40.1`. Restoring the `v0.1.31` values on the **same** engine binary:
| Engine | default | `STRATA_STAGER_THREADS=4` `STRATA_STAGER_RING=16` | `+ STRATA_PREFILL_RING=384` |
|---|---|---|---|
| `v0.1.37` | 53.9 | 355.9 | **412.8** |
| `v0.1.40.1` | 57.8 | — | **350.3** |
`v0.1.37` with the three variables reaches `v0.1.31` level (412.8 against 390-427 tok/s), 128k in 5.0 min instead of 38.3 min.
Ruled out as the cause, each measured on `v0.1.37` with the stager fix in place:
- `--pcie-frac 0.29` (the value `v0.1.31`'s probe derived): 355.9 — no change.
- `STRATA_HC_SPLIT=0` (staged hyper-connection off): 360.3 — no change.
## Proposed patch (not built, not measured — a suggestion)
The heuristic should ask where the bytes come from, not only whether the source is transient. `g_pinned_share` is already available at that point in the file:
```diff
- const int threads = stv ? std::clamp(std::atoi(stv), 1, 32) : files ? 32 : std::max(2, std::min(4, hw / 4));
- if (files && std::getenv("STRATA_STAGER_RING") == nullptr) m.stager->kRing = 4 * threads;
+ // 32 threads and a deep ring pay off when the blobs are page faults on the SSD. A RAM-resident layout
+ // answers transient() as well, but its copies come from pinned memory, where the old defaults measured
+ // 4x faster (128k prompt: 412.8 against 107.6 tok/s on an RTX 5090 / 376 GiB RAM).
+ const bool ssd = files && g_pinned_share < 0.9;
+ const int threads = stv ? std::clamp(std::atoi(stv), 1, 32) : ssd ? 32 : std::max(2, std::min(4, hw / 4));
+ if (ssd && std::getenv("STRATA_STAGER_RING") == nullptr) m.stager->kRing = 4 * threads;
```
This keeps the UD-Q4_K_XL in-place path untouched for the machines it was written for. We have not compiled or benchmarked it — the numbers above are from the existing environment variables, which is what we run in production today.
## Secondary finding
With the workaround applied, `v0.1.40.1` still runs about 15% behind `v0.1.37` (350.3 against 412.8 tok/s) on this model and prompt. That gap is independent of the stager and is a separate, much smaller regression somewhere between `v0.1.37` and `v0.1.40.1`; we have not bisected it.
## How to reproduce
1. Q8_0 pack + native shards, `--resident-budget-gib 240 --expert-cache 1944 --prefill auto --kv int8 --kv-resident 32768`.
2. Preload the model files into the page cache (`vmtouch -t <files>`) so the measurement does not include SSD reads.
3. Send a ~124k-token prompt with a fixed seed, `max_tokens: 1`, and read `prompt eval count` / `prompt eval duration` from the usage block.
4. Compare `v0.1.31` with any engine from `v0.1.32` up; then re-run the newer engine with `STRATA_STAGER_THREADS=4 STRATA_STAGER_RING=16 STRATA_PREFILL_RING=384`.
[stager-transient-fix.patch](https://github.com/user-attachments/files/33151911/stager-transient-fix.patch)En el sitio
Enlaces a install, modelos, releases.