Issues / #1349

#1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stack

open · @zyYuc · 0 comments · View on GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindows

Description

## Machine

| | |
|---|---|
| CPU / RAM | Ryzen 7 5700X, 48 GB DDR4 (4 sticks, single rank, mixed Kingbank/WODPOSIT) |
| Memory | DDR4-2666 → **DDR4-3600** via BIOS DOCP (XMP #1 = 3603 MT/s, 1.35 V, 18-22-22), FCLK 1800 (1:1) |
| GPUs | 2× RTX 2080 Ti 22 GB (sm_75), PCIe Gen3 **x8/x8** (X570 split), **NVLink NV2** (2×25.78 GB/s) |
| Engine | v0.1.40.1 (commit `82f46a8`), CUDA 13.0 source build |
| Model | Swift-1.5 Qwen3.8-Flash-Next GSQ-RCO **IQ3_XXS**, vision enabled |

Method for all tables: `run_context_ttft_fair.py` (streaming, seed = word_count×100+run), word-counts 2000/5000/8000/20000/60000 (≈2.1K/5.2K/8.3K/20.8K/62.4K tokens), 2 runs per tier, serial single stream. Engine log confirms **0 reused** on every run (true cold starts). Engine-side numbers read from the `prompt N tokens = … read in T ms` / `N generated in T ms` log lines.

## Fastest config we run (mixed workload, 262K context)

```
--resident-experts --prefill auto --spec 4 --spec-min-p 0.70 --kv int8
--max-context 262144 --vision --vram-reserve-mib 700 --pcie-frac 0.35
--ple-row-cache 8388608 --ple-inflight 1024
--pipeline-windows 2 --conversation-cache-mib 8192 --conversation-cache-slots 4
--layer-split 24
env: STRATA_SPLIT_OWN=0  STRATA_STAGE_TRIM=1  STRATA_PREFILL_HELP=1  STRATA_BF16_TC=1
```

## 1. Memory bandwidth A/B: DDR4-2666 → 3600 (+35 % theoretical)

| Tier | API prefill 2666 → 3600 | Engine prefill 2666 → 3600 |
|---|---|---|
| 2K (cold) | 470.8/563.5 → 499.1/643.4 | 503.4/610.5 → 521.2/707.1 |
| 5K | 806.4/837.4 → 834.8/874.5 | 852.6/879.0 → 878.8/913.0 |
| 8K | 880.6/924.1 → 916.8/943.4 | 910.2/953.6 → 946.2/976.9 |
| 20K | 1245.7/1293.3 → 1382.4/1368.2 | 1275.2/1325.2 → 1418.9/1403.9 |
| 60K | 1624.6/1626.9 → 1716.2/1693.1 | 1658.1/1659.7 → **1751.0/1728.1** |

**+35 % RAM bandwidth buys +3…11 % prefill** (20K best at +11 %, 60K +5 %). Decode unchanged (99 %+ expert-cache hit; MTP acceptance dominates).

`nvidia-smi dmon -s t` during a fresh 63K prompt (engine 1370 tok/s): GPU0 rx peak **7033 MB/s ≈ 90 % of the Gen3 x8 ceiling**, DMA-active only 21 of 63 s; GPU1 rx peak 5461 MB/s. The 60K window is **burst-at-the-wall + ~2/3 compute/serial** (MMQ path + CPU gather + chunk barriers): RAM bandwidth is the supply side, the PCIe upload is the pipe.

## 2. Prompt path by context size: own+4096 vs borrow+8192 vs peer-device

Engine prefill, 2-run mean, all at DDR4-2666 (peer = `--peer-device 1 --peer-reserve-mib 2048 --peer-prefill-rows 65536 --mtp-window 4096 --pcie-frac 0`, no layer-split):

| Tier | own+4096 | borrow+8192 | peer-device |
|---|---|---|---|
| 2K | 593 | 557 | **1010** |
| 5K | 818 | 866 | **1240** |
| 8K | **1175** | 923 | **1253** |
| 20K | **1394** | 1338 | **1417** |
| 60K | 1500 | **1663** | 1324 |

- **peer wins every tier ≤20K** (2K +70 %, TTFT 4.5 s → 2.4 s) and **loses 60K by −20 %** on NV2 (half the 3090 NVLink). `--peer-prefill-rows 0` collapses prefill (2K 376, −27 % vs borrow) → prompt splitting is the peer speed source, not the expert tier.
- **own+4096 best ≤20K, borrow+8192 best at 60K (+10 %)**; the 8K dip (borrow 923 vs own 1175) and the 5K borrow lead have no single explanation we found.
- Decode identical across all three modes (74–80 tok/s engine).
- own mode: `0 blob reads from the file` on all 10 requests. borrow: blob count grows monotonically (3,028 after 2K → 146,006 after 60K, ~250 GB cumulative, avg 1.7 MB/blob) — absorbed by page cache, no decay observed in this run.
- peer: 18291/24576 experts resident on the GPUs + 32.67 GiB page-locked RAM complement, log says `no file reads`; main card left with 412 MiB.

## 3. The three-switch stack (vs borrow baseline, engine-side)

`--pipeline-windows 2` + `STRATA_STAGE_TRIM=1` + `STRATA_PREFILL_HELP=1` + explicit `--layer-split 24`:

- decode, greedy 512, five content types: **75.9 → 89.3 tok/s (+17.6 %)**; windows-2 alone +11.1 %, copy-heavy +22.5 %, trim adds +5.9 %. Matches the official +16 % on the same model.
- fair 5-tier decode: 75.2 → 84.1 (+11.8 %).
- prefill: 2K 558 → 759 (+36 %, PREFILL_HELP), 5K 869 → 975, 8K 923 → 1003, 20K 1338 → 1395, 60K 1665 → 1706.

## 4. Gotchas we hit (worth documenting)

1. **server.py injects `--layer-split auto` on dual-GPU configs**, so `--peer-device` is refused. Passing `--layer-split ""` in args blocks the injection and the engine treats it as unset (`generate.cpp:1874/2228`).
2. **`STRATA_STAGE_TRIM` only trims CUDA0 under an explicit `--layer-split`** (auto only trims CUDA1, `generate.cpp:2537`) — pipeline-windows effectively forces the explicit split.
3. **`STRATA_PREFILL_HELP` prints no log line** and only applies to MMQ-path prompts ≤3.3K tokens (`prefill.cpp:1242`) — invisible when it works.
4. `./setup.sh --calibrate` picked `--pcie-frac 0.18` (+4.8 % on its own short-prompt decode sweep); in the real 5-tier bench it is **±1 % vs 0.35** — we kept 0.35.
5. 48 GB RAM: page cache (~40 GB) hugs the ~40 GB expert working set with zero headroom; capacity, not speed, is the RAM lever on this box.
6. IQ3_XXS + `STRATA_PF_FUSED`: per the v0.1.36 notes the IQ3 sizes are "about even", so never enabled (sm_75 can't run it anyway).

## Questions for maintainers

1. At 60K the non-DMA fraction is ~2/3 of the window (dmon above). Is there a knob that attacks that fraction on sm_75 — `--prefill 4096` (smaller chunks, ring streaming starts earlier), `STRATA_GR_DOWN_MAX4`, or a planned MMQ-path improvement for Turing? Happy to A/B and report.
2. Is the peer-device 60K regression on NV2 (−20 % vs borrow, independent of `--peer-prefill-rows`) expected?
3. The own/borrow crossover (own best ≤20K, borrow best 60K, 8K dip) — is the chunk-boundary interaction understood, and is a per-context-size auto-selection planned?

Raw JSONs and the full engine-log excerpts available on request. Complements #1143 (same cards): adds the memory-bandwidth variable, the prompt-path-by-size comparison, and the perf sweep of the three-switch stack.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.