Pull requests / #1391

#1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)

open · @routhjim · 0 commentaires · Sur GitHub

BenchmarksServer & APIAMD / HIPModels & quantsDocumentationWindowsLinux

Description

`STRATA_PF_FUSED=1` only reaches chunks of `stream_all_min()` tokens or more: `fused_nat` / `fused_l` in `Prefill` require `stream_all`, and `stream_all` requires `T >= stream_all_min()` (1,024). A smaller read runs its experts through MMQ. The 1,024 floor was measured on a 5070 with Q2_0, where a smaller chunk does not pay for streaming every expert over PCIe.

On Strix Halo with `--mmap-experts --expert-cache auto` every expert is resident, so nothing streams at any chunk size, and the only difference left is the kernel. There MMQ is about 3x slower than the fused kernels on a ~1,000-token read. That is the size of an agent's turn: a tool result or a test's output on top of a cached conversation. With the floor at 1,024, a third of those turns land just under it.

The switch already exists (`STRATA_PREFILL_STREAM_MIN`, the A/B override). This PR changes no code. It adds the switch to section 5's fast configuration in `docs/STRIX_HALO.md`, with a paragraph on what it does.

**It changes bits.** A read of 128-1,023 tokens gets the fused path's output, the output a read of 1,024 tokens or more already gets. Hence section 5, not the section-4 default table. It is not under `STRATA_PF_SWITCH_MIN_T`, which gates only the hyper-connection and padding switches; the fused experts already run on reads of 1,024-4,095 tokens in the fast configuration.

**Measured** on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, 128 GB, Linux 7.2, ROCm 7.14.1, the iGPU alone).
- Build: e8ca9af (main) per section 2 of `docs/STRIX_HALO.md`. Section-4 defaults, section 5's fast configuration, plus `STRATA_PF_FUSED_KQ=1` (the pack is K-quant).
- Model and flags: a C2T8 requant of Qwen3.8-Flash-Next (Q4_K / Q5_K gate/up, Q5_1 / Q8_0 down, the UD-Q4_K_XL expert formats), `--mmap-experts --expert-cache auto --prefill 16384 --no-prefill-borrow --spec 2 --kv int8`.
- Arms: one binary, two boots, `STRATA_PREFILL_STREAM_MIN=1024` (today) against `=128`.
- Times from the `strata serve: prompt ...` lines and `STRATA_TRACE=1`.

Single reads, greedy, a fresh prompt each:

| prompt read | 1024 (today) | 128 |
|---|---:|---:|
| 441 tokens | 3,107 ms | **892 ms** |
| 855 tokens | 4,039 ms | **1,227 ms** |
| 1,273 tokens (same path in both) | 1,696 ms | 1,636 ms |

An agent-shaped replay: one conversation, a 4.5K-token first prompt, then 11 turns that each append 920-1,130 tokens of source text and ask for 128 tokens. The prefix is reused every turn.

| | 1024 (today) | 128 |
|---|---:|---:|
| turns that read < 1,024 tokens (4 of 11) | 5,106-5,965 ms | **1,846-2,466 ms** |
| turns that read > 1,060 tokens (7 of 11) | 2,072-2,219 ms | 1,894-2,292 ms |
| warm turns, mean | 3,403 ms | **2,079 ms (-39%)** |

**Output.** The two reads under 1,024 tokens gave different greedy text between the arms. The 1,273-token read, on the same path in both arms, gave the same text, so the difference comes from the kernels.

**Distribution check** (teacher-forced, the method of `docs/UNSLOTH_Q4.md` / `bench/results/2026-10-03-v100-prompt-attn`).
- Sixteen chats. A first message of llama.cpp source text is read by the batched prompt path: 8 under 1,024 tokens ("short", 466-999), 8 over ("long", 1,339-2,309).
- Then a short reply and a ~270-token message that goes through the verify windows (`--short-read 400`), where every token is scored: `STRATA_LOGPOS` with `STRATA_LOGPOS_TOPK=256`, ~2,350 positions per set.
- KL is first || second over the first's top 256 plus a bucket for the rest.
- Three arms, the same binary:
  - A: today (`STREAM_MIN=1024`, `PF_FUSED_KQ=1`): short reads on MMQ, long reads fused;
  - B: this PR (128): fused for both;
  - C: a control with `PF_FUSED_KQ=0`: MMQ for both.

| comparison | reads | KL mean | median | p99 | argmax same | top-10 overlap |
|---|---|---:|---:|---:|---:|---:|
| A vs B (this PR) | short | 0.016 | 0.0011 | 0.25 | 97.1% | 94.9% |
| A vs B | long | 0 | 0 | 0 | 100% | 100% |
| A vs C (same path, a second run) | short | 0.004 | 0 | 0.10 | 99.2% | 99.3% |
| A vs C: fused vs MMQ, what the fast configuration does today | long | 0.019 | 0.0022 | 0.24 | 96.4% | 94.0% |
| C vs B: fused vs MMQ | short / long | 0.020 / 0.021 | 0.0021 / 0.0022 | 0.34 / 0.23 | 96.3 / 96.4% | 94.2 / 94.0% |

Reading:
- On short reads this PR moves the distribution as much as the fast configuration already moves every read of 1,024 tokens or more: fused vs MMQ is 0.020-0.021 at both lengths.
- Long reads do not change at all.
- The same-path short row (0.004, median 0) is the run-to-run spread of the MMQ path.

**Not tested:**
- reads between 128 and 440 tokens (the smallest batched read here was 441);
- the default `--prefill-borrow` (I measured only `--no-prefill-borrow`, which this box can afford);
- IQ packs (UD-IQ4_XS): the same gate applies, but I measured only the K-quant formats;
- Windows, and other gfx11 parts.

Measured with the help of Claude (Anthropic).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_013XBY17SSRKGmDXc2YFsrmp

Sur le site

Liens install, modèles, releases.