Issues / #1349
#1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stack
open · @zyYuc · 0 コメント · GitHub で見る
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindows
本文
## Machine | | | |---|---| | CPU / RAM | Ryzen 7 5700X, 48 GB DDR4 (4 sticks, single rank, mixed Kingbank/WODPOSIT) | | Memory | DDR4-2666 → **DDR4-3600** via BIOS DOCP (XMP #1 = 3603 MT/s, 1.35 V, 18-22-22), FCLK 1800 (1:1) | | GPUs | 2× RTX 2080 Ti 22 GB (sm_75), PCIe Gen3 **x8/x8** (X570 split), **NVLink NV2** (2×25.78 GB/s) | | Engine | v0.1.40.1 (commit `82f46a8`), CUDA 13.0 source build | | Model | Swift-1.5 Qwen3.8-Flash-Next GSQ-RCO **IQ3_XXS**, vision enabled | Method for all tables: `run_context_ttft_fair.py` (streaming, seed = word_count×100+run), word-counts 2000/5000/8000/20000/60000 (≈2.1K/5.2K/8.3K/20.8K/62.4K tokens), 2 runs per tier, serial single stream. Engine log confirms **0 reused** on every run (true cold starts). Engine-side numbers read from the `prompt N tokens = … read in T ms` / `N generated in T ms` log lines. ## Fastest config we run (mixed workload, 262K context) ``` --resident-experts --prefill auto --spec 4 --spec-min-p 0.70 --kv int8 --max-context 262144 --vision --vram-reserve-mib 700 --pcie-frac 0.35 --ple-row-cache 8388608 --ple-inflight 1024 --pipeline-windows 2 --conversation-cache-mib 8192 --conversation-cache-slots 4 --layer-split 24 env: STRATA_SPLIT_OWN=0 STRATA_STAGE_TRIM=1 STRATA_PREFILL_HELP=1 STRATA_BF16_TC=1 ``` ## 1. Memory bandwidth A/B: DDR4-2666 → 3600 (+35 % theoretical) | Tier | API prefill 2666 → 3600 | Engine prefill 2666 → 3600 | |---|---|---| | 2K (cold) | 470.8/563.5 → 499.1/643.4 | 503.4/610.5 → 521.2/707.1 | | 5K | 806.4/837.4 → 834.8/874.5 | 852.6/879.0 → 878.8/913.0 | | 8K | 880.6/924.1 → 916.8/943.4 | 910.2/953.6 → 946.2/976.9 | | 20K | 1245.7/1293.3 → 1382.4/1368.2 | 1275.2/1325.2 → 1418.9/1403.9 | | 60K | 1624.6/1626.9 → 1716.2/1693.1 | 1658.1/1659.7 → **1751.0/1728.1** | **+35 % RAM bandwidth buys +3…11 % prefill** (20K best at +11 %, 60K +5 %). Decode unchanged (99 %+ expert-cache hit; MTP acceptance dominates). `nvidia-smi dmon -s t` during a fresh 63K prompt (engine 1370 tok/s): GPU0 rx peak **7033 MB/s ≈ 90 % of the Gen3 x8 ceiling**, DMA-active only 21 of 63 s; GPU1 rx peak 5461 MB/s. The 60K window is **burst-at-the-wall + ~2/3 compute/serial** (MMQ path + CPU gather + chunk barriers): RAM bandwidth is the supply side, the PCIe upload is the pipe. ## 2. Prompt path by context size: own+4096 vs borrow+8192 vs peer-device Engine prefill, 2-run mean, all at DDR4-2666 (peer = `--peer-device 1 --peer-reserve-mib 2048 --peer-prefill-rows 65536 --mtp-window 4096 --pcie-frac 0`, no layer-split): | Tier | own+4096 | borrow+8192 | peer-device | |---|---|---|---| | 2K | 593 | 557 | **1010** | | 5K | 818 | 866 | **1240** | | 8K | **1175** | 923 | **1253** | | 20K | **1394** | 1338 | **1417** | | 60K | 1500 | **1663** | 1324 | - **peer wins every tier ≤20K** (2K +70 %, TTFT 4.5 s → 2.4 s) and **loses 60K by −20 %** on NV2 (half the 3090 NVLink). `--peer-prefill-rows 0` collapses prefill (2K 376, −27 % vs borrow) → prompt splitting is the peer speed source, not the expert tier. - **own+4096 best ≤20K, borrow+8192 best at 60K (+10 %)**; the 8K dip (borrow 923 vs own 1175) and the 5K borrow lead have no single explanation we found. - Decode identical across all three modes (74–80 tok/s engine). - own mode: `0 blob reads from the file` on all 10 requests. borrow: blob count grows monotonically (3,028 after 2K → 146,006 after 60K, ~250 GB cumulative, avg 1.7 MB/blob) — absorbed by page cache, no decay observed in this run. - peer: 18291/24576 experts resident on the GPUs + 32.67 GiB page-locked RAM complement, log says `no file reads`; main card left with 412 MiB. ## 3. The three-switch stack (vs borrow baseline, engine-side) `--pipeline-windows 2` + `STRATA_STAGE_TRIM=1` + `STRATA_PREFILL_HELP=1` + explicit `--layer-split 24`: - decode, greedy 512, five content types: **75.9 → 89.3 tok/s (+17.6 %)**; windows-2 alone +11.1 %, copy-heavy +22.5 %, trim adds +5.9 %. Matches the official +16 % on the same model. - fair 5-tier decode: 75.2 → 84.1 (+11.8 %). - prefill: 2K 558 → 759 (+36 %, PREFILL_HELP), 5K 869 → 975, 8K 923 → 1003, 20K 1338 → 1395, 60K 1665 → 1706. ## 4. Gotchas we hit (worth documenting) 1. **server.py injects `--layer-split auto` on dual-GPU configs**, so `--peer-device` is refused. Passing `--layer-split ""` in args blocks the injection and the engine treats it as unset (`generate.cpp:1874/2228`). 2. **`STRATA_STAGE_TRIM` only trims CUDA0 under an explicit `--layer-split`** (auto only trims CUDA1, `generate.cpp:2537`) — pipeline-windows effectively forces the explicit split. 3. **`STRATA_PREFILL_HELP` prints no log line** and only applies to MMQ-path prompts ≤3.3K tokens (`prefill.cpp:1242`) — invisible when it works. 4. `./setup.sh --calibrate` picked `--pcie-frac 0.18` (+4.8 % on its own short-prompt decode sweep); in the real 5-tier bench it is **±1 % vs 0.35** — we kept 0.35. 5. 48 GB RAM: page cache (~40 GB) hugs the ~40 GB expert working set with zero headroom; capacity, not speed, is the RAM lever on this box. 6. IQ3_XXS + `STRATA_PF_FUSED`: per the v0.1.36 notes the IQ3 sizes are "about even", so never enabled (sm_75 can't run it anyway). ## Questions for maintainers 1. At 60K the non-DMA fraction is ~2/3 of the window (dmon above). Is there a knob that attacks that fraction on sm_75 — `--prefill 4096` (smaller chunks, ring streaming starts earlier), `STRATA_GR_DOWN_MAX4`, or a planned MMQ-path improvement for Turing? Happy to A/B and report. 2. Is the peer-device 60K regression on NV2 (−20 % vs borrow, independent of `--peer-prefill-rows`) expected? 3. The own/borrow crossover (own best ≤20K, borrow best 60K, 8K dip) — is the chunk-boundary interaction understood, and is a per-context-size auto-selection planned? Raw JSONs and the full engine-log excerpts available on request. Complements #1143 (same cards): adds the memory-bandwidth variable, the prompt-path-by-size comparison, and the perf sweep of the three-switch stack.
関連リンク
インストール・モデル・リリースへの站内リンク。