Issues / #1586
#1586 Tester report (2x RTX 3090, 0.1.41, IQ3_S): --trim-stage-weights, CPU share on a layer split, batch 2 with and without --kv-resident
open · @adambenhassen · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows
描述
Not a bug report: measurements from a 2-GPU layer split on 0.1.41, in case they help with defaults. Everything below passed (no stalls, no watchdog restarts, answers correct).
## Setup
- 2x RTX 3090 24 GB (GPU0 on CPU PCIe 4.0 x16, GPU1 on a chipset x4 link), Ryzen 7 9800X3D (AVX-512), 32 GB RAM (30 GiB usable), Ubuntu 26.04, driver 610.43.02, No NVLink.
- Strata 0.1.41 (`fb58e0db`), engine built from source in a CUDA 13.0 container, sm_86
- Model for all runs below: Swift 1.5 IQ3_S (GSQ-RCO). Same config family also runs the original Qwen3.8-Flash-Next IQ3_S here.
- Base config: `--resident-experts`, `"layer_split": "29"`, `--pipeline-windows 2 --adapt-async 1`, `--kv int8`, `--max-context 262144`, `--spec 4 --spec-min-p 0.5`, `--vram-reserve-mib 700`, `--vision`, env `STRATA_ADAPT_LAG=2 STRATA_PF_FUSED=1` (and `STRATA_EXCHANGE_ROTATE=1`, which logs "exchange rotation unavailable" with this layout)
## Method
Each config: engine restarted, one tiny warm-up request, then before every measured request wait until both GPUs are at 52 C or below. Server `timings` fields.
- Decode: 6 x 1,500-token coding prompt, effort low, temp 0.6, seeds 1-6, mean of runs 2-6
- Short prompt reads: 3 unique prompts each of ~537 / ~983 / ~1,932 tokens, median `prompt_ms`
- Concurrency: 1 / 2 / 4 different prompts at once, 800 tokens each, total tok/s = all tokens / wall
- Mixed: a 1,500-token request, plus a short request sent 3 s later (wall time of each)
- Long context: 2 x cold ~75K-token prompt with a 1,000-token answer, then a cached follow-up turn (~850 new tokens, 500-token answer)
## 1. CPU share on a layer split (`STRATA_PREFILL_CPU_SHARE=auto`, `STRATA_PREFILL_CPU_SHARE_MAX=3072`)
Default-off on a layer split, so set explicitly. Decode unchanged, short prompts faster:
| Config | Decode tok/s | Prompt read 537 / 983 / 1,932 tokens (ms) |
|---|---:|---|
| base (two runs) | 138.0 / 138.6 | 932 / 1,358 / 2,023 |
| + CPU share | 136.8 | 870 / 1,139 / 1,676 (-7 / -16 / -17 %) |
With resident (page-locked) experts on both stages. This may be worth enabling by default on layer splits too.
## 2. `--trim-stage-weights` without `--batch`
CPU share on in both rows:
| Config | Decode tok/s | Prompt 537 / 983 / 1,932 (ms) | GPU expert cache | Pinned RAM | Free RAM |
|---|---:|---|---:|---:|---:|
| base | 137.6 | 814 / 1,094 / 1,654 | 8,230 | 16.40 GiB | 8.0 GiB |
| + `--trim-stage-weights` | **142.4** | 812 / 1,078 / 1,598 | **8,977** | **12.90 GiB** | **11.8 GiB** |
+3.5 % decode and 3.5 GiB less pinned RAM on a 32 GB machine. Setup does not write it (it needs an explicit split; setup writes `"layer_split": "auto"`).
## 3. Batch 2 on the layer split, with and without `--kv-resident`
All with trim and CPU share. The batch rows run without pipelined windows ("not with --batch slots"):
| Config | Decode alone tok/s | Prompt 537 / 983 / 1,932 (ms) | 1 / 2 / 4 at once, total tok/s | Mixed: short / long request (s) | GPU expert cache / pinned |
|---|---:|---|---|---|---|
| no batch | 142.4 | 812 / 1,078 / 1,598 | 137 / 125 / 119 | 9.6-10.2 / 10.5-11.0 | 8,977 / 12.90 GiB |
| no batch + `--kv-resident 32768` | 142.6 | 856 / 1,123 / 1,653 | 133 / 124 / 120 | 9.2-9.7 / 10.8-10.9 | 9,830 / 9.95 GiB |
| `--batch 2` | 119.8 | 1,104 / 1,543 / 2,271 | 109 / 95 / 94 | 3.5-4.1 / 15-16 | not recorded |
| `--batch 2 --kv-resident 32768` | **140.3** | 881 / 1,214 / 1,604 | 126 / 110 / 109 | **3.1-4.2** / 13.4-14.1 | 9,342 / 11.58 GiB |
Long context (~75K tokens):
| Config | Cold prompt read tok/s | Decode tok/s | Cached follow-up: ~850 new tokens read (ms) |
|---|---:|---:|---:|
| no batch | 3,247-3,284 | 100-103 | 1,018-1,062 |
| no batch + `--kv-resident 32768` | 3,207-3,235 | 103-104 | 1,161-1,193 |
| `--batch 2 --kv-resident 32768` | 2,249-2,394 | 103-108 | 1,833-1,889 |
Plain `--batch 2` costs 16 % single-request decode here. Adding `--kv-resident` gives that VRAM back to the expert cache and brings it to 140 tok/s, so a second user no longer queues behind a long request. The remaining costs are a -29 % cold long-prompt read and ~0.8 s more per cached follow-up. An earlier `--batch 4 --kv-resident 32768` run only fit 7.81 GiB pinned on this RAM and was slower everywhere (116 tok/s alone). Related to #1253 (batch MTP on multi-GPU); with MTP in batch slots, batch 2 might cost nothing here.
## 4. Smaller notes
- `STRATA_RESIDENT_HEADROOM_GIB=6` (setup's tip for 31 GB): no change here (139.8 tok/s; the 16.40 GiB complement fits either way, same free RAM).
- Re-running setup on 0.1.41 rewrote the config without `--pipeline-windows 2 --adapt-async 1` and with `"layer_split": "auto"` (log: "engine options it had that this one has not (setup chooses those)"). We restored our config by hand. Keeping options that the user set (or at least the pipelined-window pair) would help.
## Suggestions
1. Setup, 2 GPUs: write an explicit split plus `--trim-stage-weights`.
2. Consider CPU share default-on for layer splits with resident experts.
3. Setup re-run: keep user-set `--pipeline-windows` / `--adapt-async` and an explicit `layer_split`.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。