Pull requests / #1640
#1640 bench: community report, 2x RTX 3080 20 GB, UD-Q4_K_XL, 2 x 262K lanes in a 512K pool, 3.0K tok/s prompt, 86 tok/s decode
open · @noon-at-cgn · 0 comments · View on GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsSecurityDocumentationWindows
Description
## Title Issue: Related: none (results-only PR, no engine change). Nearby: #1598, #1599, #1600, #1601, #1614, #1636, #1637 (my open PRs, listed below), #1190 (resident split lend regions, by another contributor), #1253 (batch MTP on a layer split, another contributor; same design as #1636), #1011 (shared KV pool, closed by the 2026-10-06 history rewrite). ## Summary I run Qwen3.8-Flash-Next at UD-Q4_K_XL (104 GB) on a used server I built for about $2,000. A 2016 Xeon and two RTX 3080 20 GB cards read a 104K-token prompt at a median 2,996 tok/s and decode at a median 86 tok/s, with two 262,144-token lanes sharing one 524,288-token KV pool. One person measuring, one machine, 2026-10-08. My build: - Board: MACHINIST X99-MR9A PRO MAX (one socket) - CPU: Intel Xeon E5-2696 v4 (2016), 22 cores / 44 threads, up to 3.7 GHz, 55 MB L3, AVX2, no AVX-512 - RAM: 4 x 32 GB DDR4-2400 (Hynix) = 128 GB - GPUs: 2 x RTX 3080 20 GB, each PCIe Gen3 x16, each limited to 220 W - Storage: Samsung 512 GB NVMe (the model lives here) and 2 x Seagate 8 TB hard disks (other containers) - OS: Proxmox VE 9.2 host. Strata runs in an LXC with 90 GiB RAM and 24 threads - Cost: about $2,000 for the whole build Community harness (`benchmark.py`, unmodified), 3 runs per length, 256-token greedy output, median (range): | Prompt tokens | Prompt tok/s | Decode tok/s | |---|---|---| | 4,096 | 718 (708-740) | 81.2 (58.6-85.9) | | 25,000 | 2,219 (2,149-2,242) | 84.0 (79.3-85.5) | | 51,000 | 2,585 (2,515-2,601) | 90.6 (89.0-91.8) | | 103,999 | 2,996 (2,945-3,016) | 86.3 (85.0-88.4) | | 127,000 | 3,080 (3,066-3,104) | 80.7 (78.1-90.8) | | two streams, aggregate (my `ab.py`, 12 rounds x 2 identical arms, n = 24) | | 89.5 (75.4-97.2) | The two screenshots in the report show peak instantaneous performance. Prefill, 3,124 tok/s at 65,536 of 104,067 tokens (peak instantaneous). Decode, 109.4 tok/s (peak instantaneous). **This does not reproduce from upstream `main` plus the open PRs.** The engine is my fork's build on the 0.1.40.3 base (`strata-w7`, sha256 `b1df9608...`). It depends on four changes that are in neither `main` nor #1598-#1601: the shared KV pool (`--kv-pool-tokens`, my open PR #1614), `--batch-mtp` on a layer split (my open PR #1636) and `--adapt-async` beside `--batch` slots (my open PR #1637, stacked on #1636), and #1190 (resident split lend regions; an open PR by evaanp, not mine, not in `main`, currently conflicting). With a layer split and prefill borrowing, the RAM copy of the lend regions that #1190 keeps is required for the "no disk reads while serving" behaviour (start line: 5396 lent slots, all kept in RAM, 15.76 GiB). I did not measure how much the numbers depend on #1190. To get the binary I measured, build branch `repro/w7` (commit `428d9910134aaaae9ffd8c8e262cec39504221e5`) on noon-at-cgn/Strata: it is upstream v0.1.40.3 plus my patch stack. Everything in it that is not in upstream `main` is listed in `repro-w7-commits.txt` in the report folder. The sha256 of my binary is from a build of the earlier ids (`ac5fc3f` on `up-merge` `9098e89`); `repro/w7` has the same source (empty `git diff` against `ac5fc3f`) but every commit has a new id, because the wrong author e-mail on my own commits is corrected. A rebuild from `repro/w7` gives the same source, not necessarily the same binary bytes. Configuration versus patch, as measured: | Piece | Kind | Effect on this box | |---|---|---| | `--prefill auto:16384`, `STRATA_PREFILL_RING=384`, `STRATA_PREFILL_EQUAL=1` | upstream options in `main` | the largest prefill gain: 25K / 51K / 104K reads +53% / +51% / +49% together (medians of my reads) | | `--ple-io direct` | upstream option | speed neutral, frees RAM: the engine's locked memory went from 26.8 GiB to 0.0 GiB, and container `MemAvailable` from 4.6-13.8 GiB to 30.8-32.7 GiB (same pre-merge binary, 2026-10-07; readings in the README) | | [#1598](https://github.com/Niko1221/Strata/pull/1598) `--aux-cpus` | my open PR | decode +10% (solo 72 -> 80, two streams 82 -> 91 tok/s); prefill unaffected | | [#1599](https://github.com/Niko1221/Strata/pull/1599) `STRATA_SPLIT_MTP_BATCH` | my open PR | prefill +3% to +8%; decode effect not resolved (about -3% pooled) | | [#1601](https://github.com/Niko1221/Strata/pull/1601) `--memory-limit-mib` | my open PR | not a speed feature | | [#1600](https://github.com/Niko1221/Strata/pull/1600) lend sizing | my open PR (waits on #1190) | not needed here (reachable only with `STRATA_SPLIT_RING`) | | [#1190](https://github.com/Niko1221/Strata/pull/1190) resident split lend regions | open PR by evaanp, not mine, not in `main` | required for the "no disk reads while serving" behaviour with a layer split and prefill borrowing; not measured separately | | [#1614](https://github.com/Niko1221/Strata/pull/1614) shared KV pool, [#1636](https://github.com/Niko1221/Strata/pull/1636) `--batch-mtp` on a split, [#1637](https://github.com/Niko1221/Strata/pull/1637) `--adapt-async` beside slots | my open PRs | in the binary; the two-stream numbers depend on them; not measured separately on this binary (measured on 0.1.41, see Newer upstream) | | the rest of my fork (77 commits not in `main`, plus 9 merge commits; 7 of the 77 are in `main` under other commit ids) | always on in the binary: the AVX2 Q8_K quantizer (also in upstream), the O(n) window plan (#1181), the flag-B fold (`a48f52c`/`1194892` on `repro/w7`, earlier ids `530afec`/`d68cf5e`; changes the captured window graph at `--pcie-frac 0.1`, `STRATA_VERIFY_FLAGB=1` restores the old wait), per-slot vision tables (#1242), the parking fixes (#1163 and `c55444d` = #1503, earlier id `460b0a7`), the serve give-way rule (#1288) | not measured separately; I did not measure the effect of the flag-B fold and of the window plan separately | All seven of my open PRs (#1598-#1601, #1614, #1636, #1637) are in this table. ## What changed - One report folder, `bench/results/2026-10-08-community-2x-rtx3080-20gb-ud-q4-k-xl-512k-pool/`: `README.md` (my build, model and exact command, method, results, the peak screenshots, 0.1.41 numbers, what did not help), `benchmark.py` (the unmodified harness from `2026-09-30-community-rtx-5090`), `run_harness.py` (adds the API-key header, since the harness has no option for one), `ab/` (my own two-stream script and its output), `data/` (results, summary, telemetry, small request records, the 0.1.41 probe output in `w10-0.1.41/`), `repro-w7-commits.txt` (every commit of `repro/w7` that is not in v0.1.40.3, with both ids), `start-command.txt`, `strata-q4xl-prod.json`, `BUILD.json`, `machine.txt`, `model-provenance.json`, `images/`, `TRIMMED.md`. - `bench/results/COMMUNITY.md`: one row for the new folder. - No engine, server or test change. ## Extra Notes **Source and patches.** All branches are on `https://github.com/noon-at-cgn/Strata`. Production source: `repro/w7` `428d991` (= `up-merge` `9098e89`, upstream v0.1.40.3 `d5ea713` merged into my stack, plus the `opt/04` commit `ac5fc3f`; the binary does not embed its commit, the match to those earlier ids is by my build notes and by time). The binary has the 0.1.40.x-based branches of the patches that are now my open PRs (the PRs are written against 0.1.41), with merge results of those branches as text merges against upstream `main` `fb58e0d`, nothing compiled: - shared KV pool, #1614 (`pr/kv-pool-141`, head `400fc50`): binary has `kv-shared-pool` `9c533de` (`b346dd6` on `repro/w7`), copy `pr/kv-pool` `3aefa11`; base `82f46a8`, 13 commits, 393 commits behind `main`, conflicts only in code that upstream already has - batch MTP on a layer split, #1636 (`pr/batch-mtp-141`, head `2c08077`; same design as open #1253 by ilumn): binary has `batch-mtp-split` `384d67e` (`0aec553`), copy `pr/batch-mtp-split` `484b49c`, with `pr/mtp-shared-draft-head` `08ef4a3`; base `82f46a8`, 3 commits, 393 commits behind `main`, merges clean - `--adapt-async` beside `--batch` slots, #1637 (`pr/adapt-async-141`, head `4dba66a`; stacked on #1636): binary has `batch-adapt` `8d0c106` (`7c011e0`), commits `347420d` (`822a0d3`) and `3e9cecb` (`c2ce009`); base `82f46a8`, 393 commits behind `main`; `347420d` applies cleanly, `3e9cecb` conflicts with upstream `13eee45` **Newer upstream: 0.1.41.** My numbers are on engine 0.1.40.3 plus my patches (v0.1.40.3, `d5ea713`). Upstream `main` is now 0.1.41 and has 128 more commits than `up-merge`. I built my stack on 0.1.41 as branch `up-141` (head `10baecb`, code head `db925c1`, notes in `docs/UP141_NOTES.md` on that branch) and ran it as `strata-w10` on the same machine with the same production config. With my own probes (`read96`, `ab.py`), not the harness, and compared with `strata-w7` from the same probes (medians of two restarts; `strata-w10` is one start): prompt reads 25K 2,048 vs 2,128.5, 51K 2,386 vs 2,553, 104K 2,706 vs 2,940 tok/s (3.8%, 6.6%, 8.0% lower; I do not know why). Decode: one stream 79.6 and 78.7 vs 79.7 and 80.4 tok/s, two streams 88.3 and 86.9 vs 90.8 and 90.7 tok/s; the two `strata-w10` values are the two identical arms of one run, so the restart-to-restart spread of up to about 5% is not covered and I do not claim a decode change. On 0.1.41, `--batch-groups auto` is on by default for a layer split with `--batch` above 1, which runs two pipelined groups of one slot where batch MTP does not draft (upstream issue #1413); my `up-141` commit `db925c1` makes `auto` resolve to 1 when `--batch-mtp` or `--adapt-async 1` is given (start log: `batch-groups auto: 1 group of 2 slots`). #1253 is similar batch-MTP-on-a-split work by another contributor; my open PR for it is #1636. More `strata-w10` measurements from later the same day (two slots, greedy `ab.py` prompts, same binary, only the named flags differ; "pipelined groups" is `strata-w10` started with the flags upstream 0.1.41 would pick, not the upstream binary), medians: two-stream decode with batch MTP 89.2 tok/s (3 restarts) against 77.3 without it (2 restarts) and 77.7 for the pipelined groups (1 restart). The pipelined groups read a long prompt beside a running decode 2.06x faster than my serial batch windows (2,508 against 1,217 tok/s). With the adaptive tier, two-stream decode is 89.2 tok/s with the asynchronous tier, 74.7 with the blocking tier and 45.2 with the tier off. The shared KV pool pins 3.1 GiB less RAM at two lanes (72.0 against 75.1 GiB), and two-stream decode is 89.2 with the pool and 87.7 without it, inside the restart-to-restart spread. **What did not help** (same day, same box): a 280 W cap (+2% prefill, decode 0), `STRATA_STAGER_SLEEP=0`, `STRATA_AUX_STAGER=0`, `STRATA_MMVQ_IL=0`, deeper PLE I/O, ring 512, layer split 24 (-3%), `--prefill auto:32768` (+6% at 104K but the RAM is too tight), pipelined windows and two-lane pipelining. **Method.** Run on the machine that hosts the engine, against the production engine that was already running (not restarted, status idle for 32 minutes before the first request). The two-stream and one-stream numbers come from `ab.py`, my script, not the harness; its two identical arms are an A/A control, and the one-stream pair differs by 4%, which is the method's noise. Prepared with an AI assistant (Claude Code) from the raw measurements; the numbers are the engine's and the client's own output, unedited. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.