Pull requests / #1140
#1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
closed · @jase100k · 0 commentaires · Sur GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsDocumentationWindows
Description
Community benchmark report for engine **0.1.40** on a two-GPU machine. Results only — no engine source changes. Folder: `bench/results/2026-10-06-community-2x-rx-7900-gre/` — the filled template from `docs/COMMUNITY_BENCHMARKS.md`, `data/` (hardware, `engine.log`, every per-run JSON, the soak's memory samples) and `scripts/` (every script needed to repeat the runs; paths are derived from the script's own location). ## Hardware and build - 2x AMD Radeon RX 7900 GRE, 16 GiB each (gfx1100), Ryzen 7 5700X3D, 62.7 GiB RAM, NixOS 26.11, kernel 7.2.8-cachyos, ROCm 7.2.3 / HIP 7.2.53211, source build of 0.1.40. - **Lopsided PCIe link** (the main limitation): one card is on a clean 16.0 GT/s x16 path and the engine probes **28.2 GB/s** (`pcie_frac 0.55`); the other settles at 8.0 GT/s x4 and probes **3.1 GB/s** (`pcie_frac 0.08`). Layers 0-26 run on the fast card, 27-47 on the slow one, so every number below is for that pair. - Model: Qwen3.8-Flash-Next Coder IQ1_M, packs 29,608,446,496 B + 28,800,138,432 B, hashes and profile/hipblaslt-tune sha256 are in the README. Fresh prompts 1,757-4,238 tokens for the interleave, 128 for the short arms, zero reused tokens. ## Arms tested 1. **`--pipeline-windows` (#859)** — 10 interleaved pairs of 500-token greedy generations, window 2 vs window 1, one server restart per request: median decode **69.3 vs 63.8 tok/s**, paired change **-7.2%** (range -17.8% to +0.7%), window 1 slower in 9 of 10 pairs. 2. **Resident RAM on a split (#848, #642)** — `--resident-budget-gib 30`: fresh decode 72.1 tok/s, inside the baseline's own 71.2-79.1 spread. **30-minute soak: survived, 239/239 prompts, 0 errors**; RAM 22,713.6 -> 24,009.0 MiB (peak), engine RSS 10,810.3 -> 11,180.1 MiB, VRAM 15.004 -> 15.054 GiB (06) and 14.314 -> 14.440 GiB (2d), expert-cache hit rate 95.9% -> 97.2%, no `no progress`/NaN/crash line. 3. **`--batch` (#776)** — `--batch 2` and `--batch 4` load and serve; the engine switches the pipeline window off itself (`--pipeline-windows 2 is off: not with --batch slots`) and takes 1.04 / 2.08 GiB of VRAM for slot sessions. **No doorbell ever rang**: every slot finished inside the timeout (worst 43.9 s), `hung_slots` is empty in both JSONs, no stall/abort/OOM line. Filling the slots did **not** raise throughput: 68.3 tok/s one at a time, **56.6** with two at once, **64.6** with four (sum of per-slot rates), while each request fell 2.4x / 4.2x slower and TTFT stacks (slot 4 waited 18.7 s). 4. **Stage weights (#880)** — `STRATA_STAGE_TRIM=0`: 7,624 resident experts instead of 9,438 (no `77% of the experts resident` line), fresh prefill **1,582.9 -> 1,103.4 tok/s (-30.3%)**, fresh decode **75.7 -> 67.6 tok/s (-10.7%)**; both sides still print their own dense-layer range. Every configuration has at least 4 runs published with median and range; the baseline (install config) was measured as the A side of all five arms and is published as such. ## Limitations - The two cards are **not on the same link** (28.2 vs 3.1 GB/s) — stated up front in the README. - **0.1.40 only.** No version-to-version claim: `--pipeline-windows 2` was already in the 0.1.39 config but only 0.1.40 acts on it, and nothing here was measured on 0.1.39. - No answer-quality check: speed and stability only (no needle bench, no tests, no vision). - Greedy is not bit-reproducible across restarts (409 then 336 tokens for the same prompt; 3 of 10 window pairs stopped at different points) — the short runs are noisier than the 500-token ones. - One machine, one quantization, one layer split, synthetic prompts up to 4,238 fresh tokens; the 262K context, vision, tool use and sampled decoding were not tested. - The model repository **revision** was not recorded (names, sizes and download time only), and the expert profile changes between sides as the server saves it. Credentials and private paths were stripped from configs, prompts and logs; every number in the README traces to a file in `data/` (see `data/summary.md`, which is regenerated from the raw JSON).
Sur le site
Liens install, modèles, releases.