Pull requests / #583

#583 Prefill chunk ring

closed · @gopinath87607 · 0 comments · View on GitHub

BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Description

## What this is

Two things about `--prefill auto` on a pack whose expert blobs are bigger than the Q2_0 one it was
tuned on:

1. **The streamed ring is a byte budget, not a slot count.** The ring was 384 slots. That number was
   measured on Q2_0, where a slot is one expert blob of 1,382,400 B — so 506 MiB. A slot is a whole
   blob, so on a pack with bigger blobs the same 384 slots are more memory than the number was tuned
   for: **975 MiB** on this rig's IQ3_S, **1,912 MiB** on Q8_0. That memory comes out of the expert
   cache — the ring and the cache are the same VRAM.

2. **The chunk and the ring are one budget, and the ring is the better buy.** `--prefill auto` walked a
   fixed list — 32768, 16384, 8192, 6144, … — and took the first that fitted, with the ring taken off
   the top before the chunk was ever considered.

## The ring, in bytes

The budget is the constant and the slot count is derived from the pack's `max_blob`, by
`floor(budget / max_blob)`. Q2_0 still resolves to exactly **384** pinned / **96** unpinned, and the
**fused rule (1024 slots, #136) is preserved the same way**, with `ring_cap()` still bounding the
result — so 0.1.36's fused default is not undone. What changes is a pack with larger blobs, which
gets fewer slots for the same bytes:

| pack | `max_blob` | unfused (506.2 MiB) | fused (1,350.0 MiB) |
|---|---|---|---|
| Q2_0 — what it was tuned on | 1,382,400 B | 384 slots = 506.2 MiB | 1024 slots = 1,350.0 MiB |
| IQ3_S — this rig | 2,663,424 B | **199** slots = 505.5 MiB | 531 → `ring_cap()`'s **512** |
| Q8_0 | 5,222,400 B | 101 slots = 503.0 MiB | 271 slots = 1,349.7 MiB |

Flooring is why the last column is under budget rather than over: the ring never spends more than
the pack it was tuned on did. The IQ3_S fused row is the cap doing its job — 531 slots is more than
`ring_cap()` allows a native pack, so the ring stays at 512 (1,361 MiB), exactly as on `main`.
**Q2_0's own numbers are arithmetic, not a measurement**: there is no Q2_0 pack on this machine, so
`384 × 1,382,400 ÷ MAXBLOB` is the whole argument for them. The IQ3_S row is measured; the Q8_0 row
is the same arithmetic (its pack needs the PLE `Q8_0` row, which is not in this PR — see "Testing").

**One caveat, stated rather than hidden.** The fused path's own target is a *slot* count — a layer's
~460 of 512 experts, so a layer's last batch does not wait for slots its first batch frees (#136, which
is where the 1024 comes from). A byte budget can undercut that on a big-blob pack: on Q8_0 the fused
ring becomes 271 slots, below a single layer's 460. That pack's fused path needs its own A/B, which I
have not run. It is also not what this PR's measurement covers — see "Testing" for which arms were
measured and which are arithmetic.

## The auto scan

`bytes_needed` is a sum of (T × positive constant) terms plus a max of such sums, so it rises
monotonically with T, and so does every test the scan applies — which makes the largest chunk that fits
a **bisection on the 256-token grid** the prompt path already works on. Seven probes against the list's
ten, at one `bytes_needed` each. The list is coarse besides: its steps are 1,024, 2,048, 3,072, 4,096,
6,144, 8,192 — so the chunk this rig settles on, 3,584, is one it could never have named, and the gap
from 6,144 to 8,192 is 2,048 tokens.

The order was wrong because chunk and ring spend the same borrowed VRAM. Measured **with the carve in
both arms** — a different rig state from the standalone table below, so the two are not comparable — on
this rig (2× RTX 3060 + 2× RTX 5060, CUDA3 lending at its 90% cap, one cold 120K prompt per row):

| chunk | ring | prefill tok/s |
|---|---|---|
| 8960 | 17 | 963 |
| 8192 | 130 | 1008 |
| 7168 | 199 (full) | 1000 |
| 5632 | 199 (full) | 915 |

A ring slot is worth ~0.53 tok/s there and a chunk token ~0.05, so the 69 slots between a full ring and
8192's 130 are worth more than the 512 chunk tokens they cost — and once the ring *is* full, a smaller
chunk buys nothing. So the scan takes **the largest chunk that still leaves the ring full**, and only a
rig where no chunk can afford one falls back to the old rule with a `kRingMin` floor. On the same rig in
that state the rule picks 7680/199 — 1018.8 and 1021.1 tok/s on two cold 120K prompts, against the old
rule's 8960/17 at 963.1, **+5.9%**.

The room is counted in **bytes**: a ring slot is `max_blob`, but a cache slot holds its own layer's blob
— 2.15 MiB against a 2.54 MiB `max_blob` here. In slots the ring looked 18% cheaper than it is, a chunk
whose ring fitted only after that discount was refused outright, and the scan stopped a step short.

**Verified on the rig, standalone.** No carve in either arm: plain `v0.1.38` against this branch, the
same flags on both (4-way `--layer-split auto`, `--expert-cache auto`, `--prefill auto:32768`,
`--kv int8 --kv-resident 32768`, context 262144), one cold 120,000-token prompt per process, two prompts
per arm:

| | `v0.1.38` | this branch |
|---|---|---|
| auto prompt chunk | 2,048 | **3,584** |
| streamed ring | 384 slots, 975 MiB | **199 slots, 505 MiB** |
| prefill tok/s (prompt a / b) | 483.4 / 483.9 | **733.4 / 726.8** |

**+51%**, and the two halves of it are the same change. The ring is the smaller half: on this pack's
2,663,424 B blob, 384 slots is 975 MiB where the byte budget buys 199 slots for 505 MiB — Q2_0's own
506 MiB, held on a pack whose blobs are 1.9× as big. Room the ring gives back is room the chunk can
spend, and both come out of the same loan, so the same borrowed slots read 3,584 tokens at a time
instead of 2,048. (`main`'s 384 is the old rule's number, and it is what that binary runs: `fused_ring()`
is false on this native pack unless the native fused kernels are opted into with `STRATA_PF_FUSED=1`, so
`STRATA_PF_FUSED=0` reproduces the default run line for line. With `=1` the same two binaries pick
1,024 → 1,536 — and there the ring is 512 slots in both, since `ring_cap()` clamps a native pack to 512,
so that arm isolates the scan on its own and the chunk still grows.)

## Also

`bytes_needed` counted `carve`'s buffers from memory and got two wrong: it took `xn` unconditionally,
though `carve` takes it only under `STRATA_GR_UNFUSED`, and it never took `grs`. Net T*(D−HC)*4 bytes
too many — 42 MB at a 1,024-token chunk, 252 MB at 6,144. Safe in direction (the prompt path was told
it had less room than it did), but it under-sizes every loan.

And the Monitor's GPU panel gains a "Prompt chunk" row reading the engine's new `prefill_chunk` /
`prefill_ring` INFO fields, so the pair the run settled on is visible from the web app. `prefill_ring` is
the **resolved** ring, not the room the scan offered: `ring_slots()` clamps to [16, RING_MAX], a chunk
under `stream_all_min()` gets STAGE, and `STRATA_PREFILL_RING` overrides both.

## Testing

`serve`'s 119 Python tests pass on `v0.1.38` (`python3 -m unittest serve.test_server`), including
`PromptChunkFacts` for the new INFO fields.

The prompt-path numbers are the four runs of the standalone A/B — two cold 120,000-token prompts per arm,
one prompt per process, the same flags on both arms, same base for both binaries — plus the carve-present
sweep, which was measured in the state this branch was developed in.

Q8_0 is the one row in the byte table that is arithmetic rather than a run here: its pack needs the PLE
`Q8_0` row, which is not in this PR, so `v0.1.38` refuses it (`per_layer_token_embd.weight is Q8_0, not
IQ4_NL, Q5_0 or FP8 (I8)`). The Q8_0 measurement the comments cite — a ring forced by `STRATA_PREFILL_RING`
on the four-way rig, chunk 1,024 → 6,144 and prefill 87 → 402 tok/s — is from the branch this was split
out of, which is the only one here that can load that pack.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.