Issues / #519
#519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt path
open · @brenoperucchi · 7 comentarios · En GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Descripción
mise ~/.config/mise/config.toml tools: [email protected] Two observations from one machine, in case they help with the default for the IQ sizes and with short prompts. **Machine:** RTX 5090 32 GB (PCIe Gen 4 x16; the engine's PCIe probe reports 28.3 GB/s host-to-device when started without `--pcie-frac`), Ryzen 9 5950X (AVX2 only), 96 GB DDR4-3200, Windows 11, driver 616.64. Engine 0.1.36 and 0.1.34 release binaries. **Model and config:** Swift IQ3_XXS with `--expert-cache auto --prefill auto:32768 --spec 4 --spec-min-p 0.70 --mtp <path to the MTP layer> --max-context 32768 --kv int8 --pcie-frac 0.20` and `STRATA_IQ_MT_MIN=1`. This is the config this PC serves in production. The expert cache holds 14,864 of the 24,576 experts (24.17 GiB). **Method:** the same as in #433: `/v1/chat/completions`, temperature 0, 256-token cap, 1 warm-up + 3 measured runs per prompt, each run reading its whole prompt fresh (`cache_n` 0). Prompts built from this repo at `v0.1.32`. Engine `timings`, medians in tokens/s. ## 1. STRATA_PF_FUSED=1 on IQ3_XXS | Engine | ~2.7K tokens | ~14.7K tokens | ~28.9K tokens | |---|---:|---:|---:| | 0.1.34 | 2,650 | 5,513 | 6,222 | | 0.1.36 | 2,690 | 5,534 | 6,231 | | 0.1.36, `STRATA_PF_FUSED=1` | 3,192 (+19%) | 6,784 (+23%) | 7,043 (+13%) | The release notes measured the IQ3 sizes about even on the RTX 5070, so on this card the result differs. Decode medians were 158-169 tok/s without the flag and 158-192 with it; the highest one (~2.7K) also accepted more drafts (143-155 vs 127-133), so we don't read it as an effect of the flag. We only measured speed: the benchmark did not keep the answer text, so we have not compared the answers with the default kernels. ## 2. Short prompts: ~470 ms in the batched prompt path Most requests this PC serves are classification calls: a 250-600 token prompt and a 1-token answer, about 80 a minute on one slot. We measured that shape with a ~555-token prompt (554-560 tokens across the runs) and `max_tokens: 1`, 2 warm-ups and 10 runs: | Config | Prompt time, median | |---|---:| | 0.1.34 | 539 ms | | 0.1.36 | 526 ms | | 0.1.36, `STRATA_PF_FUSED=1` | 529 ms | | 0.1.36, `STRATA_PF_FUSED=1`, `--prefill auto` (8K chunks) | 532 ms | That is ~1,050 tokens/s, against ~3,200 for the ~2.7K-token prompt on the same server. A separate run with `STRATA_TRACE=1` (6 requests of 549-550 tokens) shows where one of them spends its time: ``` strata trace: request 549 0 strata trace: prompt start 548 -1 strata trace: lent 265 slots for 768 tokens in 0.1 ms strata trace: prompt chunk 0 of 544 strata trace: read 544 tokens (batched) in 470.1 ms strata trace: refill start 265 -1 strata trace: refilled 265 slots on 1 stage(s) in 18.8 ms strata trace: read 4 tokens (windows) in 14.5 ms strata trace: prompt done (slots refilled) -1 -1 strata trace: window 548 1 strata serve: prompt 549 tokens = 0 reused + 549 read in 529 ms (1037.2 tok/s), 1 generated in 9 ms (115.8 tok/s), drafts accepted 0 of 0, 1 checkpoints ``` The batched read took 468-474 ms in the five traced requests after the first (which took 619 ms); the loan and the refill take ~19 ms, and the 4-token header goes through the windows. With `STRATA_PF_FUSED=1` set, the whole prompt takes 529 ms for ~556 tokens and 839 ms for 2,679 tokens: about 5x the tokens for 1.6x the time. (The short prompt is under `stream_all_min()`'s 1,024 tokens, so it does not use the fused kernels; the two points come from different paths, and a line through them is only a rough estimate of the part that does not grow with length.) The comment above `windows_ok` in `src/program/generate.cpp` (v0.1.36) puts the batched path at ~300 ms per run however few tokens it has, plus ~180 ms to refill the borrowed slots. Here the split is different: ~470 ms in the batched read and ~19 ms in the refill, with about 40% of the experts outside VRAM. Question: is that split what you would expect on this card, and is there anything besides `--short-read` (default 64, from the engine's `--help`) to lower it for prompts of a few hundred tokens? The windows took 14.5-18.6 ms per 4 tokens here (3.6-4.7 ms per token), so by a linear estimate a larger `--short-read` would only pay under ~100-130 tokens. Next on our side: these classification prompts may share a fixed instruction, so we will try a checkpoint at its end (`--prompt-cache-root` below its 2,048-token default, from `docs/DETAILS.md`) with the varying part under `--short-read`, and report what it does to the per-request time.
En el sitio
Enlaces a install, modelos, releases.