Issues / #781

#781 1M context: identical expert work per window, but the VRAM-touching stages grow 7-8x

open · @1314521gjy · 7 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsWindows

本文

Windows, RTX 4080 SUPER (32 GB, driver 616.56, PCIe Gen 4 x16), i7-13790F (8P+16E, AVX-2), 96 GB RAM. Strata **0.1.38**, Flash-Next **IQ3_S**, `--kv int8 --kv-resident 32768 --pool-workers 4 --spec 4 --pcie-frac 0.35`. Everything is identical between the three runs below except `--max-context` and the rope settings it requires (262,144 / 524,288 + yarn 2 / 1,048,576 + yarn 4). The expert cache is **11,631 slots / 22.07 GiB in all three** (bundled profile, no eviction).

Same prompt, 18 decode windows, 64 tokens, with `STRATA_DECODE_TIMING=1`, `STRATA_VERIFY_PROFILE=1`, `STRATA_PREFILL_TIMING=1` (ms per window):

| stage | 262,144 | 524,288 | 1,048,576 |
| --- | ---: | ---: | ---: |
| ms/window | 38.47 | 35.30 | **69.15** |
| GPU-reach wait | 12.94 | 12.33 | **41.52** |
| GDN VRAM hits | 3.61 | 3.58 | **25.35** |
| QSA VRAM hits | 1.23 | 1.23 | **10.27** |
| GDN PCIe grp | 0.28 | 0.28 | **1.81** |
| q8 gemv / attention / out-proj | 1.26 / 0.74 / 0.93 | 1.27 / 0.73 / 0.89 | 1.37 / 0.78 / 1.00 |
| work: experts/layer, VRAM hits/window | 2.51, 33.24 | 2.58, 34.22 | 2.54, 33.74 |

On the prompt side, the same prompt read cold: GPU timeline 2,392 -> 2,273 -> 3,880 ms, where `qsa proj` is 159 -> 145 -> 306 ms while `hc read` / `gdn` (314/312 -> 298/290) barely move.

End to end, fresh 43,969-token read + 200 generated: 262,144 = 2,581 ms read (2,652 tok/s) / 103.9 tok/s decode; 524,288 = 2,655 ms / 94.4; 1,048,576 = 4,065 ms (1,684 tok/s) / 46.4.

**Reading:** the work counters do not change and the expert cache is identical, yet the stages that touch VRAM cost 7-8x more and the host's wait for the GPU grows 3.4x, while pure-compute stages move about 10%. That looks like access cost, not extra work. 524,288 shows none of it (its stage table is within noise of 262,144).

**Hypothesis, not measured:** the per-layer session tables scale with `--max-context` (KV residency page table, QSA indexer tables, GDN state), so the same expert addresses may become more expensive to reach (TLB pressure, or a 4x larger session state in VRAM breaking locality). I did not measure TLB or page-fault counters and did not measure the VRAM carve under a fixed `--vram-reserve-mib`.

**Question:** is the 1M tier expected to cost this much by design, or is there a knob that recovers it (a different `--kv-resident` window, `--pcie-frac`, or separate residency sizing for the KV page table at very large `--max-context`)?

YaRN is not the cause: at 1,048,576, yarn factor 4 was faster than no scaling at both sample sizes (46.4 vs 39.3 tok/s; 44.2 vs 41.4 on a 64-token sample), with the same expert cache and the same 32,768 resident KV cells. Happy to re-run any of this with extra instrumentation on the same machine.

関連リンク

インストール・モデル・リリースへの站内リンク。