反馈 / #1683

#1683 [Feature]: --resident-budget-gib keeps no lend region, so every prompt chunk reads the borrowed slots' experts from the SSD

open · @KarlGrier · 0 评论 · 去 GitHub 看

BenchmarksNVIDIA / CUDAModels & quantsWindows

说明

### What do you want to do?

Use `--resident-budget-gib` (to cap how much RAM the expert copy takes) without the prompt path reading its borrowed cache slots' experts from the SSD on every chunk, when a larger budget, or RAM, has room for them.

Today a RAM budget turns the RAM copy's lend region off. `pin_cache_complement` sets `lend_from_slot = -1` in its budget branch ([expert_source.cpp L2291 at v0.1.41](https://github.com/Niko1221/Strata/blob/fb58e0dbc8399662c0e47c76578c6e878b14f6cf/src/core/expert_source.cpp#L2291), commented "A budget turns `lend` off"). So the experts of the GPU slots that a prompt chunk borrows are in neither VRAM nor RAM during the prompt, and every chunk reads them from the drive, even when the budget already holds every other expert and RAM has room for these too. A larger budget does not help: the budget branch only places experts the GPU cache does not hold. Without a budget, `--resident-experts` keeps them ("2225 of the prompt path's 2225 lendable slots keep their experts in RAM too").

**Measured 2026-10-09 on stock v0.1.41** (source build with `-DSTRATA_MMQ_KQUANTS=ON`). UD-Q4_K_XL (unsloth, 71.7 GiB of routed experts), RTX 5080 16 GB alone, 96 GB RAM, `--prefill auto:16384`. Served, A B B A, a fresh server per arm, model page cache dropped first. Two different 32,768-token prompts per arm, nothing reused:

| arm | RAM copy | lent slots' experts in RAM | 32K prompt read (s) | prompt tok/s | read from the drive per 32K prompt | decode tok/s | MemAvailable, lowest |
|---|---|---|---|---|---|---|---|
| `--resident-budget-gib 66` | 64.00 GiB (every expert the GPU cache does not hold) | 0 of 2,225 | 11.53-12.03 | 2,724-2,844 | 28.8-29.2 GB | 61.0-68.4 | 16.6 GiB |
| `--resident-experts` | 70.49 GiB | 2,225 of 2,225 | 9.14-9.50 | 3,449-3,584 | 1.2 GB | 66.6-71.1 | 12.2 GiB |

Four 32K reads per setting (two arms, two prompts each). The budget arm takes about 26 % longer to read a 32K prompt (11.76 s against 9.33 s on average) and reads about 24 times as much from the drive. The engine's own I/O line for one 32K prompt: "the OS read 28821.4 MB from storage (14 major faults) for 27507.6 MB of expert reads" with the budget, and "the OS read 1215.3 MB from storage ... for 0.0 MB of expert reads" with `--resident-experts`. Decode after these prompts was 61.0-68.4 tok/s with the budget and 66.6-71.1 with `--resident-experts`. These are short answers, one per prompt; we did not look into that difference.

The start warning in budget mode says:

```
strata generate: WARNING: a prompt chunk of 16384 tokens lends 2225 cache slots to the prompt path, but only 0 of their experts fit in RAM: the others are read from the SSD on every chunk, and prompts can read ~3x slower than with --prefill auto (#669). A smaller --prefill, or auto, keeps them all in RAM
```

With a budget the count is 0 whatever the chunk size, so its advice (a smaller `--prefill`, or auto) cannot keep them in RAM here.

Two ways this could go: the budget mode keeps a lend region inside the budget, filled after the complement as far as the budget reaches (as `--resident-experts` does from available RAM). Or the warning names the budget as the cause when one is set.

Related: #669 and #765 (the warning itself), #1389 (lend coverage with `--resident-experts`, where the line now warns when it is not N of N), and #1190 (the same gap for a layer split; it leaves the single-GPU path as it is).

Not tested: a helper GPU, a layer split, Windows, other files and chunk sizes, `--prefill auto`, and longer prompts (on 2026-10-08 a local build of v0.1.41 showed the same warning in budget mode and read a 120K prompt in 34 s).

Measured and drafted with Claude Code.

本站相关内容

相关页面的快捷入口。