Issues / #519

#519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt path

open · @brenoperucchi · 7 コメント · GitHub で見る

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

本文

mise ~/.config/mise/config.toml tools: [email protected]
Two observations from one machine, in case they help with the default for the IQ sizes and with short prompts.

**Machine:** RTX 5090 32 GB (PCIe Gen 4 x16; the engine's PCIe probe reports 28.3 GB/s host-to-device when started
without `--pcie-frac`), Ryzen 9 5950X (AVX2 only), 96 GB DDR4-3200, Windows 11, driver 616.64. Engine 0.1.36 and
0.1.34 release binaries.
**Model and config:** Swift IQ3_XXS with `--expert-cache auto --prefill auto:32768 --spec 4 --spec-min-p 0.70
--mtp <path to the MTP layer> --max-context 32768 --kv int8 --pcie-frac 0.20` and `STRATA_IQ_MT_MIN=1`. This is the
config this PC serves in production. The expert cache holds 14,864 of the 24,576 experts (24.17 GiB).
**Method:** the same as in #433: `/v1/chat/completions`, temperature 0, 256-token cap, 1 warm-up + 3 measured runs
per prompt, each run reading its whole prompt fresh (`cache_n` 0). Prompts built from this repo at `v0.1.32`.
Engine `timings`, medians in tokens/s.

## 1. STRATA_PF_FUSED=1 on IQ3_XXS

| Engine | ~2.7K tokens | ~14.7K tokens | ~28.9K tokens |
|---|---:|---:|---:|
| 0.1.34 | 2,650 | 5,513 | 6,222 |
| 0.1.36 | 2,690 | 5,534 | 6,231 |
| 0.1.36, `STRATA_PF_FUSED=1` | 3,192 (+19%) | 6,784 (+23%) | 7,043 (+13%) |

The release notes measured the IQ3 sizes about even on the RTX 5070, so on this card the result differs. Decode
medians were 158-169 tok/s without the flag and 158-192 with it; the highest one (~2.7K) also accepted more drafts
(143-155 vs 127-133), so we don't read it as an effect of the flag. We only measured
speed: the benchmark did not keep the answer text, so we have not compared the answers with the default kernels.

## 2. Short prompts: ~470 ms in the batched prompt path

Most requests this PC serves are classification calls: a 250-600 token prompt and a 1-token answer, about 80 a
minute on one slot. We measured that shape with a ~555-token prompt (554-560 tokens across the runs) and
`max_tokens: 1`, 2 warm-ups and 10 runs:

| Config | Prompt time, median |
|---|---:|
| 0.1.34 | 539 ms |
| 0.1.36 | 526 ms |
| 0.1.36, `STRATA_PF_FUSED=1` | 529 ms |
| 0.1.36, `STRATA_PF_FUSED=1`, `--prefill auto` (8K chunks) | 532 ms |

That is ~1,050 tokens/s, against ~3,200 for the ~2.7K-token prompt on the same server. A separate run with
`STRATA_TRACE=1` (6 requests of 549-550 tokens) shows where one of them spends its time:

```
strata trace: request 549 0
strata trace: prompt start 548 -1
strata trace: lent 265 slots for 768 tokens in 0.1 ms
strata trace: prompt chunk 0 of 544
strata trace: read 544 tokens (batched) in 470.1 ms
strata trace: refill start 265 -1
strata trace: refilled 265 slots on 1 stage(s) in 18.8 ms
strata trace: read 4 tokens (windows) in 14.5 ms
strata trace: prompt done (slots refilled) -1 -1
strata trace: window 548 1
strata serve: prompt 549 tokens = 0 reused + 549 read in 529 ms (1037.2 tok/s), 1 generated in 9 ms (115.8 tok/s), drafts accepted 0 of 0, 1 checkpoints
```

The batched read took 468-474 ms in the five traced requests after the first (which took 619 ms); the loan and the
refill take ~19 ms, and the 4-token header goes through the windows. With `STRATA_PF_FUSED=1` set, the whole prompt
takes 529 ms for ~556 tokens and 839 ms for 2,679 tokens: about 5x the tokens for 1.6x the time. (The short prompt is
under `stream_all_min()`'s 1,024 tokens, so it does not use the fused kernels; the two points come from different
paths, and a line through them is only a rough estimate of the part that does not grow with length.)

The comment above `windows_ok` in `src/program/generate.cpp` (v0.1.36) puts the batched path at ~300 ms per run
however few tokens it has, plus ~180 ms to refill the borrowed slots. Here the split is different: ~470 ms in the
batched read and ~19 ms in the refill, with about 40% of the experts outside VRAM.

Question: is that split what you would expect on this card, and is there anything besides `--short-read` (default
64, from the engine's `--help`) to lower it for prompts of a few hundred tokens? The windows took 14.5-18.6 ms per 4
tokens here (3.6-4.7 ms per token), so by a linear estimate a larger `--short-read` would only pay under ~100-130
tokens.

Next on our side: these classification prompts may share a fixed instruction, so we will try a checkpoint at its end
(`--prompt-cache-root` below its 2,048-token default, from `docs/DETAILS.md`) with the varying part under
`--short-read`, and report what it does to the per-request time.

関連リンク

インストール・モデル・リリースへの站内リンク。