Issues / #1280
#1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
open · @andrej-reimer · 1 comentários · No GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
**2x RTX 3090, Linux: `--pipeline-windows 2` is a win in every expert mode (+7% to +20% decode), and a 30-minute soak of the resident RAM mode on the split ran clean. Two observations below that may be worth a look.** The 0.1.40 notes ask for 2-GPU testers: a 30-minute soak of the resident RAM mode on a split (#848, #642) and `--pipeline-windows` with 10 or more interleaved pairs, with and without #848 (#859). This covers both on a 2x RTX 3090 pair. I could not test #776: IQ3_S at 262K does not fit fully into 2x 24 GB here, so neither stage is all-resident. ## Headline Median decode tok/s over 5 greedy prompts, engine timings, fresh server per run, interleaved ABBA: | Expert mode | pairs | serial | `--pipeline-windows 2` | | prompt 31.5K, both arms | |---|---:|---:|---:|---:|---:| | `--resident-experts` (#848) | 10 | 89.4 (86.8-90.9) | **107.4** (102.7-109.5) | **+20%** | 17.1 / 17.2 s | | `--mmap-experts` | 10 | 107.1 (83.8-109.3) | **114.5** (110.4-116.8) | +7% | 18.5 / 18.6 s | | no flag (whole arena in RAM) | 5 | 110.4 (110.2-110.8) | **126.0** (125.3-126.5) | +14% | 14.8 / 14.8 s | In the resident arms the ranges do not overlap: the slowest pipelined run beat the fastest serial one in all 10 pairs. ## Rig - Proxmox VE host, AMD Ryzen 9 7950X, 2x 48 GB DDR5-5200 (dual channel). The engine runs in an Ubuntu 24.04.4 VM (kernel 6.8.0-142) with both GPUs passed through: 16 vCPUs, **60 GB RAM**. - **GPU 0:** RTX 3090 (reference-design board), PCIe 4.0 **x16** (on a riser). **GPU 1:** ASUS ROG Strix RTX 3090, PCIe 4.0 **x4**. `nvidia-smi topo -m`: PHB, no P2P, no NVLink. - Both cards power-capped at **250 W**. - Driver 595.91.07. Strata **v0.1.40** (`1735d64`), built with the repo's Dockerfile (`CUDA_ARCHITECTURES=86`, nvcc 13.0), so `serve/` and the engine are a matched pair. ## Model and settings - Qwen3.8-Flash-Next **IQ3_S** (GSQ-RCO pack from setup), MTP draft layer, vision on. - `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 262144 --kv int8 --vision --vram-reserve-mib 700`, `"layer_split": "auto"`, `STRATA_POOL_SPIN_US=0`. - The engine chose `K=22 - predicted 25.6 ms per decode window; the caches hold 16476 of 24576 profiled pairs (~98.8% of the routed mass)`. - `"remote_expert_opt": false` in every arm, since #859 turns itself off with it. I did not test what the default does on a split. - The three arms differ only in the expert flag (`--mmap-experts`, `--resident-experts`, or neither) plus `--pipeline-windows 2`. The pipeline logged `two verifiers per stage, the stages overlap across windows`; the resident arm logged `resident RAM mode: 16.35 GiB of experts in RAM (page-locked), 8823 in the GPU cache`. ## Method - My own harness (not the upstream tools). Each run recreates the container, sends one warmup, then 5 prompts at `temperature: 0`, `enable_thinking: false`, `max_tokens: 500`: - a Python function (ISO-8601 duration parser) - a C linked-list reversal - an English short story - a German essay - a copy-heavy edit: return a 1.5 KB Python file with one function renamed - Then one 31.5K-token prompt (the Strata README with a unique prefix, `max_tokens: 1`). - Decode tok/s and prompt time come from `/metrics`. Per prompt (median tok/s, serial → pipelined): | | Python | C | English prose | German prose | edit | |---|---:|---:|---:|---:|---:| | resident | 93.5 → 115.0 | 107.2 → 134.7 | 67.7 → 75.8 | 71.0 → 76.8 | 104.8 → 133.8 | | mmap | 112.8 → 120.2 | 127.9 → 142.8 | 79.7 → 80.4 | 82.4 → 80.3 | 130.1 → 147.3 | | whole arena | 118.3 → 133.3 | 135.1 → 161.1 | 83.6 → 86.2 | 87.7 → 87.5 | 127.3 → 160.5 | The gain follows code and copy-heavy text, as in #859's own table. ## Resident RAM mode soak (#848) - **30 minutes, 346 requests** at `temperature: 0.6`, thinking on for half of them, `max_tokens` 300/600/1000, eight prompts (code, English, German, French). - **0 degenerate or empty answers** (n-gram repetition check plus non-empty). No engine exit, no container restart. - Decode mean 84.7 tok/s (first 10 requests 79.6, last 10 94.8). - Engine counter at the end: `15.94 GiB of experts in RAM, 523349 exchanged with the VRAM tier, 81342 blob reads from the file`. - Memory of the container at the end: shmem 17.5 GiB (the locked copy), anon 2.4 GiB, file cache 35.4 GiB. ## Two observations **1. Blob reads from the file in the resident mode.** The start line says the swaps happen with no file reads, but the `blob reads from the file` counter grows, mostly during long prompts. Across one 31.5K-token prompt it went from 777 to 18,777, and it reached 81,342 over the soak. In this 60 GB VM the page cache serves those reads, so the prompt time was unaffected. On a 32 GB PC, the case #848 targets, they would hit the SSD. I have not looked into which path issues them; this is an observation, not a diagnosis. **2. Resident is slower than mmap here, but this is not the low-RAM case.** With 60 GB of RAM, the OS file cache holds the whole 46.8 GiB `experts.bin`, so `--mmap-experts` effectively runs from RAM, and it beat `--resident-experts` (107 vs 89 serial, 114 vs 107 pipelined). I did not test a 32 GB VM, where I would expect the opposite. Happy to run that if it helps. ## Exactness I recorded output hashes only for the mmap and whole-arena arms (I added them after the resident pairs had run): - **Whole arena, serial:** byte-identical across all 5 runs, on every prompt. - **Whole arena, pipelined:** - C and the edit are identical to serial. - Python had 2 variants in 5 runs (differing in its first line: `import re` placed before the function or not). - The two prose prompts gave a different text in every run. Examples: "the rhythmic sweep of the **Fresnel lens**" vs "of the **beam**"; „öffentliche Wohnzimmer“ vs „Wohnzimmer der Stadt“; some runs open with a different first sentence. - **mmap, serial** also varies (3-6 variants per prompt in 10 runs; C once `exit(1)` instead of `exit(EXIT_FAILURE)`), so not all of the variation comes from the pipeline. - All 59 distinct outputs are fluent and correct. None repeats or is empty. The prose cuts off at my 500-token limit, which is my test design, not the model. I did not run with `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`, so I make no claim about bit-exactness with equal caches. This only says that with the default adaptive tier, the pipelined prose is not reproducible run to run on this rig while the serial whole-arena run is. ## Not claimed - No power numbers for these runs. - No 0.1.39 vs 0.1.40 comparison. Our 0.1.39 numbers were taken with a different harness. - Only one model and quant (IQ3_S). - The rig is a VM with passthrough, not bare metal. Happy to rerun any arm, with the exactness flags, or at 32 GB.
No site
Links install, modelos, releases.