Issues / #1009
#1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)
open · @hiru0118 · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
本文
*Posted on behalf of a friend.* The measurements, investigation and write-up below are theirs. They could not post it from their own environment, so I am posting it for them. If you have questions, I will check with them and reply here.
---
On an 8 GB card, 0.1.39's default `--prefill auto` reads long prompts about 40% slower than 0.1.31 did on the same PC: 190-197 tok/s became 114-117 tok/s on a 22K-token prompt. The log shows the prompt chunk at 256 tokens where 0.1.31 used 512. Adding `--prefill 512 --no-prefill-borrow` brings it back (208 tok/s).
### Environment
- GPU: RTX 4060 Ti 8 GB (PCIe 4.0 x8), driver 610.88, Windows 11
- CPU: Core i5-14400F (no AVX-512), RAM: 32 GB DDR4-3200, SSD: Kingston NV2 1 TB (NVMe)
- Strata: v0.1.39 (6f32ec0), ready-made engine 0.1.39 (CUDA 13.0). Compared with 0.1.31 (9259cad) on the same PC.
- Model: Coder IQ1_M. Setup: `--family coder --model IQ1_M --context 32768 --kv int8 --vision no --experimental-speed-projection off`
- Config args as setup wrote them (paths shortened):
```
--pack <data>/packs/coder-iq1_m --native <data>/models/coder-IQ1_M/...-00001-of-00002.gguf
--ple-gguf <data>/models/coder-IQ1_M/...-00002-of-00002.gguf --expert-profile <strata>/data/expert-profile-coder.bin
--expert-cache auto --prefill auto --spec 4 --mtp <data>/mtp/rt --max-context 32768 --kv int8
--resident-experts --spec-min-p 0.5
```
I know 8 GB is below the README's 12 GB recommendation. I am reporting it because 0.1.31 was faster on this card with the same config.
### Measurements
Each row is one run. Prefill and decode are the server's `timings`; TTFT is to the first token (thinking included).
| engine / args | expert cache | prompt chunk | prefill 8K | prefill 22K | TTFT 22K | decode 22K |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| 0.1.31, default (run 1) | 181-240 slots | 512 | 148.5 | 189.9 | 118.3 s | 19.1 |
| 0.1.31, default (run 2) | 181-240 slots | 512 | 147.2 | 196.6 | 114.2 s | 20.0 |
| 0.1.39, default | 309 slots | 256 | 113.7 | 116.9 | 191.9 s | 19.0 |
| 0.1.39, default + `STRATA_RING_BYTES=0` | 314 slots | 256 | 113.8 | 113.9 | 196.8 s | 19.5 |
| 0.1.39, `--prefill 512 --no-prefill-borrow` | 1 slot | 512 (see below) | 193.4 | 207.7 | 108.1 s | 21.2 |
Prompts were 6,851-6,853 and 22,370-22,375 tokens. On 0.1.31 the cache size varied between starts (181-240 slots).
### Logs
0.1.31:
```
strata serve: the prompt path allocates its own buffers (too few cache slots to borrow)
strata serve: a 1024-token chunk's prompt buffers need 830 MiB on CUDA0, 921 MiB free: 512-token chunks
```
0.1.39, default:
```
strata generate: expert cache 309 slots, 0.60 GiB of VRAM; policy is
strata serve: prompt chunk auto: 256 tokens, a 8-slot ring
strata serve: the prompt path borrows 146 CUDA0 cache slots (0.28 GiB)
strata serve: prompt 22375 tokens = 0 reused + 22375 read in 191355 ms (116.9 tok/s),
404 generated in 21226 ms (19.0 tok/s), drafts accepted 263 of 372, 2 checkpoints
```
0.1.39 with only `--prefill 512` (start-up only, not benchmarked):
```
strata generate: expert cache 309 slots, 0.60 GiB of VRAM; policy is
strata serve: prompt chunk 512 -> 256 tokens so its buffers fit in every expert cache
strata serve: the prompt path borrows 146 CUDA0 cache slots (0.28 GiB)
```
0.1.39 with `--prefill 512 --no-prefill-borrow`:
```
strata generate: expert cache auto: 1.25 GiB free, 700 MiB reserved (+218 MiB for the draft head) -> 0 slots
strata generate: expert cache auto: the 700 MiB reserve leaves too few slots on this card
(a working cache needs 1): a 554 MiB reserve instead -> 1 slots
strata serve: the prompt path allocates its own buffers (too few cache slots to borrow)
strata serve: prompt 22375 tokens = 0 reused + 22375 read in 107740 ms (207.7 tok/s),
401 generated in 18872 ms (21.2 tok/s), drafts accepted 263 of 355, 2 checkpoints
```
The log does not print the chunk for this last run. I read it as 512 because I asked for 512, there is no "512 -> 256" line, and prefill came back to the 0.1.31 level.
### Why, as far as I can tell
A 256-token chunk borrows 146 of the 309 cache slots, so a 512-token chunk would need roughly twice that, about 290. The loan has to leave 128 slots in the cache (`fits_one` in `src/program/generate.cpp`), and 290 + 128 > 309, so the chunk drops to 256. `STRATA_RING_BYTES=0` does not change this (prefill stayed at 114 tok/s).
### Workaround
Add `--prefill 512 --no-prefill-borrow` to the config's args. On this PC the expert cache then shrinks to 1 slot (the decode hit rate went from 14-22% to 0.1%), but decode did not get slower (21.2 vs 19.0 tok/s on the 22K prompt). Setup does not carry hand-added args over when it rewrites the config, so this has to be re-added after each update.
### Caveats
- One run per config. Between the two 0.1.31 runs, prefill differed by 1-4% and decode by up to 20%.
- The amount of experts held in RAM by `--resident-experts` was not the same across runs (0.1.31: none, it fell back to mmap; 0.1.39 default: 15.7 GiB; `STRATA_RING_BYTES=0`: 21.0 GiB; the workaround: 18.5 GiB), because free RAM at start differed. Prefill did not follow it (113.9-116.9 tok/s at 15.7 and 21.0 GiB).
- Not tested on a 12 GB or larger card. I expect a larger cache leaves room for a bigger borrowed chunk there.
### Possible change
When owned buffers allow a larger chunk than the borrowing plan can get, the auto choice could take the owned path, as on 0.1.31. #658 proposes something close to this (measured on an RTX 5090). Related: #765, #796.
### How it was measured
For each config the server was restarted, then one warm-up request, a 56-token code request, an ~8K prompt plus one follow-up, and a ~22K prompt plus one follow-up, all through `/v1/chat/completions` with streaming, `reasoning_effort: medium`, `temperature: 0`. The long prompts are Strata's own source files from 9259cad followed by a short question, so they are nearly the same token count on both engines. The OS file cache was not controlled.
関連リンク
インストール・モデル・リリースへの站内リンク。