Issues / #1373
#1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it
closed · @Scorp1o117 · 5 commentaires · Sur GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Description
--- ### Resolution (final): the original report was a cold-measurement artifact — with a warmed engine, 0.1.40.2 is **12.9% faster** here, not slower **The mistake was mine, in the benchmark.** My harness restarted the engine for every data point and took the decode measurement after only three requests. Strata's expert cache adapts while the server runs — it replaces cold experts with the ones actually being routed — so a freshly loaded engine is measurably slower than the same engine a few minutes later. I was measuring the warm-up ramp, not the engine. Measured properly — **one load, ten decodes inside the same process**, same config, same server layer, same frozen prompt, and the two engines run back to back in one evening: | engine | decode, 448 tokens, ten runs | mean | 13.8K prompt | 41K prompt | |---|---|---:|---:|---:| | 0.1.39 | 121.3 131.6 129.6 126.0 128.7 126.9 127.2 126.0 123.9 133.8 | 127.5 | 3,430.8 | 3,838.3 | | **0.1.40.2** | 135.1 139.3 145.7 144.2 142.1 147.1 146.9 145.7 143.8 149.9 | **144.0** | 3,312.9 | 3,851.6 | So 0.1.40.2 is **+12.9% decode** — the two ranges barely overlap and the within-run spread is about 10% — in the same direction as the +2.5 to +6% you measured on the RTX 5070. Nothing is broken. The earlier "update" that said the two were indistinguishable was wrong for the same reason. **Please ignore this whole report**, including the `STRATA_MMVQ_IL` isolation, the "−20%" and the free-VRAM observations: they all came from the same broken method. Two things in it were method-independent and still hold: - `--no-capture` is not usable here: the engine refuses it (`--no-capture runs 'session_token', which has NO CPU expert pool hook, so the routed experts would silently contribute nothing`). - `--vram-reserve-mib 4096` on this config moves the auto split to K=41, drops the cache to 13,278 slots and the hit rate from 88% to 81%. The lesson, in case it saves someone else the same detour: **warm up first, then measure.** Restart the engine as few times as possible and take repeated decodes inside one process. With that method I also found that a preceding 41K-token prompt does *not* degrade the decode that follows (134.0 and 133.2 right after it — the fastest numbers of the run), and that the within-version spread is about 10%, so anything smaller than that needs paired runs to be believed. Closing this — sorry for the noise. I am still glad to run the two-GPU work you asked for (resident RAM mode on a split, `--batch` on an all-resident stage, `--pipeline-windows`, with and without #848), this time with the warm-up method above. --- ### Update, same evening (~2 h later): this does not survive more runs — please read this part first I ran **0.1.39 four more times tonight** on the same machine, same session, same frozen prompt, and it is **just as slow as 0.1.40.2**: | engine | decode, 448 tokens, tonight's runs | mean | |---|---|---| | 0.1.40.2 | 111.9, 104.8 | **108.4** | | 0.1.39 | 74.8, 126.9, 105.4, 124.5 | **107.9** | The same 0.1.39 binary that measured 135.0 / 135.5 / 135.6 two days ago now spans **74.8 to 126.9 (70%)** within one evening. So the −20% in the report below is **an artifact of comparing against a two-day-old baseline**, not a 0.1.40.2 regression. Sorry for the noise — I would rather correct it now than let you spend time on it. What still holds, and what does not: - **The two versions are indistinguishable on this box** at this level of noise (mean decode 107.9 vs 108.4). - **`STRATA_MMVQ_IL` is unconfirmed too.** The pair I measured (104.8 → 124.6, five minutes apart) sits inside the band 0.1.39 itself produces tonight (74.8–126.9), so it cannot be separated from interference. - **The machine is roughly 25% slower tonight than on 2026-10-05, with much higher run-to-run variance**, and I do not yet know what changed on it. Every cross-session comparison in this report inherits that. - Still accurate: the new `strata verify: capturing the N-token window` lines, the free-VRAM drop after them, `--no-capture` being refused by the engine, and `--vram-reserve-mib 4096` moving the auto split to K=41 with the cache at 13,278 slots and the hit rate at 81%. The right experiment for a difference this small is **paired, interleaved runs** — engine A and B alternating across 4-6 pairs in one session, which cancels exactly this kind of drift. I am happy to run that on this machine and post the pairs, in either direction. The two-GPU offer in the last section stands. ### What happened I upgraded one installation from 0.1.39 to 0.1.40.2 (engine + Python layer; the `strata-*.json` was not touched). On the same machine, same model and same config, **every** measurement got slower: | | 0.1.39 (3 runs, 2026-10-05) | 0.1.40.2 (2 runs) | change | |---|---:|---:|---:| | 13.8K-token prompt, prefill | 3,263.7 tok/s | 2,742.9 tok/s | **−16.0%** | | 41K-token prompt, prefill | 3,735.4 tok/s | 3,588.6 tok/s | −3.9% | | 448-token answer, decode | 135.4 tok/s | 108.3 tok/s | **−20.0%** | I bisected it. **`STRATA_MMVQ_IL=0` recovers 15% of the decode** — the new "a 2-4 token verify window reads each weight block once", on by default on RTX 30 and newer: | | 13.8K prompt | 41K prompt | decode | |---|---:|---:|---:| | 0.1.39 | 3,263.7 | 3,735.4 | **135.4** | | 0.1.40.2, defaults | 2,742.9 | 3,588.6 | 108.3 | | 0.1.40.2 + `STRATA_MMVQ_IL=0` | 3,028.4 | 3,653.8 | 124.6 | | 0.1.40.2 + `STRATA_MMVQ_IL=0` + `--vram-reserve-mib 4096` | 3,023.0 | 3,414.0 | 126.7 | So here the new kernel costs 15% of the decode instead of the +1.3 to +2.2% you measured on one RTX 5070, and with it off this box is still 8% below 0.1.39. ### Strata version, GPU, OS - **0.1.40.2** (engine `strata.exe` sha256 `04a4df567870…`, 159,099,904 B), CUDA 13.0 build - Windows 11 (build 26300), NVIDIA driver 617.14 - GPU 0: **RTX 4080 SUPER 32 GB** (sm_89, 80 SMs @ 2.55 GHz) · GPU 1: **RTX 5070 Ti 16 GB** (sm_120, 70 SMs @ 2.48 GHz) - AMD Ryzen Threadripper 3970X (32C/64T), 128 GB RAM - Model: **Swift 1.5 IQ3_XXS** native Strata pack (70.6 GiB over two shards; the weights are a community abliteration of Swift 1.5, but the pack layout is Swift's and the model is constant across every run below) - Args: `--pack … --native (shard 1) --ple-gguf (shard 1) --expert-profile … --expert-cache auto --prefill auto --spec 4 --mtp … --max-context 262144 --kv int8 --vision --spec-min-p 0.70 --kv-resident 32768 --pool-workers 15 --pool-affinity all --pcie-frac 0.20 --vram-reserve-mib 2048`, `"gpu": [0, 1]`, `"layer_split": "auto"` → **K=39** ### How I measured - **The prompt is frozen**: `docs/DETAILS.md` exactly as shipped in 0.1.39 (90,707 B, sha256 `85EDA834…`), so the 0.1.39 and 0.1.40.2 prompt tok/s read the same text. The 41K prompt is that document three times (80,494 tokens, 16,384 of them reused from the conversation cache, so 64,110 read). Token counts differ by 3 and 9 tokens between the versions, from the tokenizer change. - The two prompt reads use `reasoning_effort: none`, `enable_thinking: false`, `max_tokens: 16`; the decode is 448 tokens of plain prose. Each run is a fresh engine start with the model reloaded, and nothing else runs on the machine. - Both engines ran through **the same server** (0.1.40.2's `serve/`), so only the engine binary differs between rows. - Every run reports the same engine-side state, so the two versions are running the same model in the same shape: K=39 split, 14,531 expert cache slots (77% of the experts resident), 1,825 MiB free with everything loaded, 98.3% decode cache hit rate. ### What I ruled out - **`--no-capture` is not a usable workaround**: the engine refuses it — `--no-capture runs 'session_token', which has NO CPU expert pool hook, so the routed experts would silently contribute nothing. Pass --no-pool as well if the GPU-only floor is what you want.` With `--no-pool` the CPU pool is gone, which is not the config I am after. - **`--vram-reserve-mib 4096`** buys 2% decode and costs 6.6% on the 41K prompt: the auto split moves to K=41, the cache drops 14,531 → 13,278 slots and the hit rate 88% → 81%. Not worth it. - **Not a slower machine**: loading is *faster* now (39.97 GiB at 3.09 / 3.80 GiB/s on 0.1.39 vs 3.91 / 3.92 GiB/s on 0.1.40.2), and the two 0.1.40.2 runs agree (2,792 / 2,693 · 3,569 / 3,608 · 111.9 / 104.8). - The 15% from `STRATA_MMVQ_IL` is a **same-session A/B** (two runs about five minutes apart, one environment variable). The remaining 7-8% is against the 0.1.39 numbers from two days earlier, so I would want a back-to-back pair to confirm that part — happy to run one. ### Possibly related, not measured 0.1.40.2 prints `strata verify: capturing the N-token window` lines that 0.1.39 does not, and the free VRAM after them is much lower here: `1825 MiB → 157-180 MiB` with the defaults, `1825 → 210-428 MiB` with `STRATA_MMVQ_IL=0`. This box is tight — the engine already shrinks the cache to the reserve (`only 537 MiB free once the slots are written (reserve 2048 MiB)`) — so if the interleaved-activation path needs VRAM, the captures may be interacting with it. That part is a guess; the measured part is the 15%. (Your #1275 notes "the verify captures are 57.2 MiB of device buffers" on a single gfx1200; here the per-window drop I read in the log is much larger, which may or may not be the same allocation.) ### Ask - Is `STRATA_MMVQ_IL` expected to help on a **two-GPU layer split**, or is the interleaved activation copy meant for a single card? If it is meant to work here, I can run any A/B you want on this box. - You asked for two-GPU testers in the 0.1.40 notes. This machine (2 GPUs, 48 GB VRAM, 128 GB RAM, Windows, native pack) is available — say the word and I will run `--pipeline-windows`, `--adapt-async`, the resident RAM mode on a split (#848) and the `--batch` cases, and post the numbers. ### Engine log 0.1.40.2, defaults, startup + the four measurements (`strata-swift-iq3_xxs.log`): ``` strata generate: layer split across 2 GPUs: CUDA0, then CUDA1 (split auto) strata generate: layer split auto: CUDA0 80 SMs at 2.55 GHz -> 0.36 ms per layer, 25.42 GiB free before its session carve strata generate: layer split auto: CUDA1 70 SMs at 2.48 GHz -> 0.42 ms per layer, 8.39 GiB free before its session carve strata generate: layer split auto: K=39 - predicted 18.7 ms per decode window; the caches hold 20013 of 24576 profiled pairs (~99.4% of the routed mass), best of 46 placements strata generate: expert cache auto: 26.68 GiB free, 2048 MiB reserved (+0 MiB for the draft head) -> 11373 slots strata generate: only 537 MiB free once the slots are written (reserve 2048 MiB); shrinking the expert cache strata generate: expert cache 14531 slots, 23.14 GiB of VRAM; policy is strata generate: pre-filled 14531 of 14531 slots from the profile in 2.2 s (11090 MB/s); slot 0 verified strata generate: layer split, CUDA1: 9.95 GiB free of 15.89, room for experts 7.85 GiB strata generate: layer split: CUDA1 runs layers 39-47, expert cache 4356 slots (7.85 GiB), 4356 of its 4608 profiled pairs; slot 0 verified strata generate: layer split: CUDA0 runs layers 0-38 strata generate: layer split auto: L2 probe: the split is startable on the live free figures strata generate: session is up (engine 0.1.40.2) strata serve: layer split: 77% of the experts resident, the prompt path's streamed ring 96 slots strata serve: the prompt path borrows 2252 CUDA0 cache slots (3.63 GiB) strata serve: CUDA1 prompt path borrows 1996 of its 4356 slots (3.63 GiB) strata serve: layer split: layers 0-38 (CUDA0), 39-47 (CUDA1), one hand-off per window strata serve: 1825 MiB of VRAM free with everything loaded strata verify: capturing the 6-token window (1825 MiB of VRAM free) strata verify: capturing the 6-token window (333 MiB of VRAM free) strata verify: capturing the 1-token window (1813 MiB of VRAM free) strata verify: capturing the 1-token window (305 MiB of VRAM free) strata verify: capturing the 4-token window (1801 MiB of VRAM free) strata v
Sur le site
Liens install, modèles, releases.