Pull requests / #1140

#1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776

closed · @jase100k · 0 comentários · No GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPModels & quantsDocumentationWindows

Descrição

Community benchmark report for engine **0.1.40** on a two-GPU machine. Results only — no engine
source changes.

Folder: `bench/results/2026-10-06-community-2x-rx-7900-gre/` — the filled template from
`docs/COMMUNITY_BENCHMARKS.md`, `data/` (hardware, `engine.log`, every per-run JSON, the soak's
memory samples) and `scripts/` (every script needed to repeat the runs; paths are derived from the
script's own location).

## Hardware and build

- 2x AMD Radeon RX 7900 GRE, 16 GiB each (gfx1100), Ryzen 7 5700X3D, 62.7 GiB RAM,
  NixOS 26.11, kernel 7.2.8-cachyos, ROCm 7.2.3 / HIP 7.2.53211, source build of 0.1.40.
- **Lopsided PCIe link** (the main limitation): one card is on a clean 16.0 GT/s x16 path and the
  engine probes **28.2 GB/s** (`pcie_frac 0.55`); the other settles at 8.0 GT/s x4 and probes
  **3.1 GB/s** (`pcie_frac 0.08`). Layers 0-26 run on the fast card, 27-47 on the slow one, so
  every number below is for that pair.
- Model: Qwen3.8-Flash-Next Coder IQ1_M, packs 29,608,446,496 B + 28,800,138,432 B, hashes and
  profile/hipblaslt-tune sha256 are in the README. Fresh prompts 1,757-4,238 tokens for the
  interleave, 128 for the short arms, zero reused tokens.

## Arms tested

1. **`--pipeline-windows` (#859)** — 10 interleaved pairs of 500-token greedy generations, window 2
   vs window 1, one server restart per request: median decode **69.3 vs 63.8 tok/s**, paired
   change **-7.2%** (range -17.8% to +0.7%), window 1 slower in 9 of 10 pairs.
2. **Resident RAM on a split (#848, #642)** — `--resident-budget-gib 30`: fresh decode 72.1 tok/s,
   inside the baseline's own 71.2-79.1 spread. **30-minute soak: survived, 239/239 prompts, 0
   errors**; RAM 22,713.6 -> 24,009.0 MiB (peak), engine RSS 10,810.3 -> 11,180.1 MiB, VRAM
   15.004 -> 15.054 GiB (06) and 14.314 -> 14.440 GiB (2d), expert-cache hit rate 95.9% -> 97.2%,
   no `no progress`/NaN/crash line.
3. **`--batch` (#776)** — `--batch 2` and `--batch 4` load and serve; the engine switches the
   pipeline window off itself (`--pipeline-windows 2 is off: not with --batch slots`) and takes
   1.04 / 2.08 GiB of VRAM for slot sessions. **No doorbell ever rang**: every slot finished inside
   the timeout (worst 43.9 s), `hung_slots` is empty in both JSONs, no stall/abort/OOM line.
   Filling the slots did **not** raise throughput: 68.3 tok/s one at a time, **56.6** with two at
   once, **64.6** with four (sum of per-slot rates), while each request fell 2.4x / 4.2x slower and
   TTFT stacks (slot 4 waited 18.7 s).
4. **Stage weights (#880)** — `STRATA_STAGE_TRIM=0`: 7,624 resident experts instead of 9,438 (no
   `77% of the experts resident` line), fresh prefill **1,582.9 -> 1,103.4 tok/s (-30.3%)**, fresh
   decode **75.7 -> 67.6 tok/s (-10.7%)**; both sides still print their own dense-layer range.

Every configuration has at least 4 runs published with median and range; the baseline (install
config) was measured as the A side of all five arms and is published as such.

## Limitations

- The two cards are **not on the same link** (28.2 vs 3.1 GB/s) — stated up front in the README.
- **0.1.40 only.** No version-to-version claim: `--pipeline-windows 2` was already in the 0.1.39
  config but only 0.1.40 acts on it, and nothing here was measured on 0.1.39.
- No answer-quality check: speed and stability only (no needle bench, no tests, no vision).
- Greedy is not bit-reproducible across restarts (409 then 336 tokens for the same prompt; 3 of 10
  window pairs stopped at different points) — the short runs are noisier than the 500-token ones.
- One machine, one quantization, one layer split, synthetic prompts up to 4,238 fresh tokens; the
  262K context, vision, tool use and sampled decoding were not tested.
- The model repository **revision** was not recorded (names, sizes and download time only), and the
  expert profile changes between sides as the server saves it.

Credentials and private paths were stripped from configs, prompts and logs; every number in the
README traces to a file in `data/` (see `data/summary.md`, which is regenerated from the raw JSON).

No site

Links install, modelos, releases.