Issues / #1373

#1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it

closed · @Scorp1o117 · 5 comentários · No GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Descrição

---

### Resolution (final): the original report was a cold-measurement artifact — with a warmed engine, 0.1.40.2 is **12.9% faster** here, not slower

**The mistake was mine, in the benchmark.** My harness restarted the engine for every data point and took the decode measurement after only three requests. Strata's expert cache adapts while the server runs — it replaces cold experts with the ones actually being routed — so a freshly loaded engine is measurably slower than the same engine a few minutes later. I was measuring the warm-up ramp, not the engine.

Measured properly — **one load, ten decodes inside the same process**, same config, same server layer, same frozen prompt, and the two engines run back to back in one evening:

| engine | decode, 448 tokens, ten runs | mean | 13.8K prompt | 41K prompt |
|---|---|---:|---:|---:|
| 0.1.39 | 121.3 131.6 129.6 126.0 128.7 126.9 127.2 126.0 123.9 133.8 | 127.5 | 3,430.8 | 3,838.3 |
| **0.1.40.2** | 135.1 139.3 145.7 144.2 142.1 147.1 146.9 145.7 143.8 149.9 | **144.0** | 3,312.9 | 3,851.6 |

So 0.1.40.2 is **+12.9% decode** — the two ranges barely overlap and the within-run spread is about 10% — in the same direction as the +2.5 to +6% you measured on the RTX 5070. Nothing is broken.

The earlier "update" that said the two were indistinguishable was wrong for the same reason. **Please ignore this whole report**, including the `STRATA_MMVQ_IL` isolation, the "−20%" and the free-VRAM observations: they all came from the same broken method.

Two things in it were method-independent and still hold:

- `--no-capture` is not usable here: the engine refuses it (`--no-capture runs 'session_token', which has NO CPU expert pool hook, so the routed experts would silently contribute nothing`).
- `--vram-reserve-mib 4096` on this config moves the auto split to K=41, drops the cache to 13,278 slots and the hit rate from 88% to 81%.

The lesson, in case it saves someone else the same detour: **warm up first, then measure.** Restart the engine as few times as possible and take repeated decodes inside one process. With that method I also found that a preceding 41K-token prompt does *not* degrade the decode that follows (134.0 and 133.2 right after it — the fastest numbers of the run), and that the within-version spread is about 10%, so anything smaller than that needs paired runs to be believed.

Closing this — sorry for the noise. I am still glad to run the two-GPU work you asked for (resident RAM mode on a split, `--batch` on an all-resident stage, `--pipeline-windows`, with and without #848), this time with the warm-up method above.

---

### Update, same evening (~2 h later): this does not survive more runs — please read this part first

I ran **0.1.39 four more times tonight** on the same machine, same session, same frozen prompt, and it is **just as slow as 0.1.40.2**:

| engine | decode, 448 tokens, tonight's runs | mean |
|---|---|---|
| 0.1.40.2 | 111.9, 104.8 | **108.4** |
| 0.1.39 | 74.8, 126.9, 105.4, 124.5 | **107.9** |

The same 0.1.39 binary that measured 135.0 / 135.5 / 135.6 two days ago now spans **74.8 to 126.9 (70%)** within one evening. So the −20% in the report below is **an artifact of comparing against a two-day-old baseline**, not a 0.1.40.2 regression. Sorry for the noise — I would rather correct it now than let you spend time on it.

What still holds, and what does not:

- **The two versions are indistinguishable on this box** at this level of noise (mean decode 107.9 vs 108.4).
- **`STRATA_MMVQ_IL` is unconfirmed too.** The pair I measured (104.8 → 124.6, five minutes apart) sits inside the band 0.1.39 itself produces tonight (74.8–126.9), so it cannot be separated from interference.
- **The machine is roughly 25% slower tonight than on 2026-10-05, with much higher run-to-run variance**, and I do not yet know what changed on it. Every cross-session comparison in this report inherits that.
- Still accurate: the new `strata verify: capturing the N-token window` lines, the free-VRAM drop after them, `--no-capture` being refused by the engine, and `--vram-reserve-mib 4096` moving the auto split to K=41 with the cache at 13,278 slots and the hit rate at 81%.

The right experiment for a difference this small is **paired, interleaved runs** — engine A and B alternating across 4-6 pairs in one session, which cancels exactly this kind of drift. I am happy to run that on this machine and post the pairs, in either direction. The two-GPU offer in the last section stands.

### What happened

I upgraded one installation from 0.1.39 to 0.1.40.2 (engine + Python layer; the `strata-*.json` was not touched). On the same machine, same model and same config, **every** measurement got slower:

| | 0.1.39 (3 runs, 2026-10-05) | 0.1.40.2 (2 runs) | change |
|---|---:|---:|---:|
| 13.8K-token prompt, prefill | 3,263.7 tok/s | 2,742.9 tok/s | **−16.0%** |
| 41K-token prompt, prefill | 3,735.4 tok/s | 3,588.6 tok/s | −3.9% |
| 448-token answer, decode | 135.4 tok/s | 108.3 tok/s | **−20.0%** |

I bisected it. **`STRATA_MMVQ_IL=0` recovers 15% of the decode** — the new "a 2-4 token verify window reads each weight block once", on by default on RTX 30 and newer:

| | 13.8K prompt | 41K prompt | decode |
|---|---:|---:|---:|
| 0.1.39 | 3,263.7 | 3,735.4 | **135.4** |
| 0.1.40.2, defaults | 2,742.9 | 3,588.6 | 108.3 |
| 0.1.40.2 + `STRATA_MMVQ_IL=0` | 3,028.4 | 3,653.8 | 124.6 |
| 0.1.40.2 + `STRATA_MMVQ_IL=0` + `--vram-reserve-mib 4096` | 3,023.0 | 3,414.0 | 126.7 |

So here the new kernel costs 15% of the decode instead of the +1.3 to +2.2% you measured on one RTX 5070, and with it off this box is still 8% below 0.1.39.

### Strata version, GPU, OS

- **0.1.40.2** (engine `strata.exe` sha256 `04a4df567870…`, 159,099,904 B), CUDA 13.0 build
- Windows 11 (build 26300), NVIDIA driver 617.14
- GPU 0: **RTX 4080 SUPER 32 GB** (sm_89, 80 SMs @ 2.55 GHz) · GPU 1: **RTX 5070 Ti 16 GB** (sm_120, 70 SMs @ 2.48 GHz)
- AMD Ryzen Threadripper 3970X (32C/64T), 128 GB RAM
- Model: **Swift 1.5 IQ3_XXS** native Strata pack (70.6 GiB over two shards; the weights are a community abliteration of Swift 1.5, but the pack layout is Swift's and the model is constant across every run below)
- Args: `--pack … --native (shard 1) --ple-gguf (shard 1) --expert-profile … --expert-cache auto --prefill auto --spec 4 --mtp … --max-context 262144 --kv int8 --vision --spec-min-p 0.70 --kv-resident 32768 --pool-workers 15 --pool-affinity all --pcie-frac 0.20 --vram-reserve-mib 2048`, `"gpu": [0, 1]`, `"layer_split": "auto"` → **K=39**

### How I measured

- **The prompt is frozen**: `docs/DETAILS.md` exactly as shipped in 0.1.39 (90,707 B, sha256 `85EDA834…`), so the 0.1.39 and 0.1.40.2 prompt tok/s read the same text. The 41K prompt is that document three times (80,494 tokens, 16,384 of them reused from the conversation cache, so 64,110 read). Token counts differ by 3 and 9 tokens between the versions, from the tokenizer change.
- The two prompt reads use `reasoning_effort: none`, `enable_thinking: false`, `max_tokens: 16`; the decode is 448 tokens of plain prose. Each run is a fresh engine start with the model reloaded, and nothing else runs on the machine.
- Both engines ran through **the same server** (0.1.40.2's `serve/`), so only the engine binary differs between rows.
- Every run reports the same engine-side state, so the two versions are running the same model in the same shape: K=39 split, 14,531 expert cache slots (77% of the experts resident), 1,825 MiB free with everything loaded, 98.3% decode cache hit rate.

### What I ruled out

- **`--no-capture` is not a usable workaround**: the engine refuses it — `--no-capture runs 'session_token', which has NO CPU expert pool hook, so the routed experts would silently contribute nothing. Pass --no-pool as well if the GPU-only floor is what you want.` With `--no-pool` the CPU pool is gone, which is not the config I am after.
- **`--vram-reserve-mib 4096`** buys 2% decode and costs 6.6% on the 41K prompt: the auto split moves to K=41, the cache drops 14,531 → 13,278 slots and the hit rate 88% → 81%. Not worth it.
- **Not a slower machine**: loading is *faster* now (39.97 GiB at 3.09 / 3.80 GiB/s on 0.1.39 vs 3.91 / 3.92 GiB/s on 0.1.40.2), and the two 0.1.40.2 runs agree (2,792 / 2,693 · 3,569 / 3,608 · 111.9 / 104.8).
- The 15% from `STRATA_MMVQ_IL` is a **same-session A/B** (two runs about five minutes apart, one environment variable). The remaining 7-8% is against the 0.1.39 numbers from two days earlier, so I would want a back-to-back pair to confirm that part — happy to run one.

### Possibly related, not measured

0.1.40.2 prints `strata verify: capturing the N-token window` lines that 0.1.39 does not, and the free VRAM after them is much lower here: `1825 MiB → 157-180 MiB` with the defaults, `1825 → 210-428 MiB` with `STRATA_MMVQ_IL=0`. This box is tight — the engine already shrinks the cache to the reserve (`only 537 MiB free once the slots are written (reserve 2048 MiB)`) — so if the interleaved-activation path needs VRAM, the captures may be interacting with it. That part is a guess; the measured part is the 15%. (Your #1275 notes "the verify captures are 57.2 MiB of device buffers" on a single gfx1200; here the per-window drop I read in the log is much larger, which may or may not be the same allocation.)

### Ask

- Is `STRATA_MMVQ_IL` expected to help on a **two-GPU layer split**, or is the interleaved activation copy meant for a single card? If it is meant to work here, I can run any A/B you want on this box.
- You asked for two-GPU testers in the 0.1.40 notes. This machine (2 GPUs, 48 GB VRAM, 128 GB RAM, Windows, native pack) is available — say the word and I will run `--pipeline-windows`, `--adapt-async`, the resident RAM mode on a split (#848) and the `--batch` cases, and post the numbers.

### Engine log

0.1.40.2, defaults, startup + the four measurements (`strata-swift-iq3_xxs.log`):

```
strata generate: layer split across 2 GPUs: CUDA0, then CUDA1 (split auto)
strata generate: layer split auto: CUDA0 80 SMs at 2.55 GHz -> 0.36 ms per layer, 25.42 GiB free before its session carve
strata generate: layer split auto: CUDA1 70 SMs at 2.48 GHz -> 0.42 ms per layer, 8.39 GiB free before its session carve
strata generate: layer split auto: K=39 - predicted 18.7 ms per decode window; the caches hold 20013 of 24576 profiled pairs (~99.4% of the routed mass), best of 46 placements
strata generate: expert cache auto: 26.68 GiB free, 2048 MiB reserved (+0 MiB for the draft head) -> 11373 slots
strata generate: only 537 MiB free once the slots are written (reserve 2048 MiB); shrinking the expert cache
strata generate: expert cache 14531 slots, 23.14 GiB of VRAM; policy is
strata generate: pre-filled 14531 of 14531 slots from the profile in 2.2 s (11090 MB/s); slot 0 verified
strata generate: layer split, CUDA1: 9.95 GiB free of 15.89, room for experts 7.85 GiB
strata generate: layer split: CUDA1 runs layers 39-47, expert cache 4356 slots (7.85 GiB), 4356 of its 4608 profiled pairs; slot 0 verified
strata generate: layer split: CUDA0 runs layers 0-38
strata generate: layer split auto: L2 probe: the split is startable on the live free figures
strata generate: session is up (engine 0.1.40.2)
strata serve: layer split: 77% of the experts resident, the prompt path's streamed ring 96 slots
strata serve: the prompt path borrows 2252 CUDA0 cache slots (3.63 GiB)
strata serve:   CUDA1 prompt path borrows 1996 of its 4356 slots (3.63 GiB)
strata serve: layer split: layers 0-38 (CUDA0), 39-47 (CUDA1), one hand-off per window
strata serve: 1825 MiB of VRAM free with everything loaded
strata verify: capturing the 6-token window (1825 MiB of VRAM free)
strata verify: capturing the 6-token window (333 MiB of VRAM free)
strata verify: capturing the 1-token window (1813 MiB of VRAM free)
strata verify: capturing the 1-token window (305 MiB of VRAM free)
strata verify: capturing the 4-token window (1801 MiB of VRAM free)
strata v

No site

Links install, modelos, releases.