Issues / #1009

#1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)

open · @hiru0118 · 0 comments · View on GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

Description

*Posted on behalf of a friend.* The measurements, investigation and write-up below are theirs. They could not post it from their own environment, so I am posting it for them. If you have questions, I will check with them and reply here.

---

On an 8 GB card, 0.1.39's default `--prefill auto` reads long prompts about 40% slower than 0.1.31 did on the same PC: 190-197 tok/s became 114-117 tok/s on a 22K-token prompt. The log shows the prompt chunk at 256 tokens where 0.1.31 used 512. Adding `--prefill 512 --no-prefill-borrow` brings it back (208 tok/s).

### Environment

- GPU: RTX 4060 Ti 8 GB (PCIe 4.0 x8), driver 610.88, Windows 11
- CPU: Core i5-14400F (no AVX-512), RAM: 32 GB DDR4-3200, SSD: Kingston NV2 1 TB (NVMe)
- Strata: v0.1.39 (6f32ec0), ready-made engine 0.1.39 (CUDA 13.0). Compared with 0.1.31 (9259cad) on the same PC.
- Model: Coder IQ1_M. Setup: `--family coder --model IQ1_M --context 32768 --kv int8 --vision no --experimental-speed-projection off`
- Config args as setup wrote them (paths shortened):

```
--pack <data>/packs/coder-iq1_m --native <data>/models/coder-IQ1_M/...-00001-of-00002.gguf
--ple-gguf <data>/models/coder-IQ1_M/...-00002-of-00002.gguf --expert-profile <strata>/data/expert-profile-coder.bin
--expert-cache auto --prefill auto --spec 4 --mtp <data>/mtp/rt --max-context 32768 --kv int8
--resident-experts --spec-min-p 0.5
```

I know 8 GB is below the README's 12 GB recommendation. I am reporting it because 0.1.31 was faster on this card with the same config.

### Measurements

Each row is one run. Prefill and decode are the server's `timings`; TTFT is to the first token (thinking included).

| engine / args | expert cache | prompt chunk | prefill 8K | prefill 22K | TTFT 22K | decode 22K |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| 0.1.31, default (run 1) | 181-240 slots | 512 | 148.5 | 189.9 | 118.3 s | 19.1 |
| 0.1.31, default (run 2) | 181-240 slots | 512 | 147.2 | 196.6 | 114.2 s | 20.0 |
| 0.1.39, default | 309 slots | 256 | 113.7 | 116.9 | 191.9 s | 19.0 |
| 0.1.39, default + `STRATA_RING_BYTES=0` | 314 slots | 256 | 113.8 | 113.9 | 196.8 s | 19.5 |
| 0.1.39, `--prefill 512 --no-prefill-borrow` | 1 slot | 512 (see below) | 193.4 | 207.7 | 108.1 s | 21.2 |

Prompts were 6,851-6,853 and 22,370-22,375 tokens. On 0.1.31 the cache size varied between starts (181-240 slots).

### Logs

0.1.31:

```
strata serve: the prompt path allocates its own buffers (too few cache slots to borrow)
strata serve: a 1024-token chunk's prompt buffers need 830 MiB on CUDA0, 921 MiB free: 512-token chunks
```

0.1.39, default:

```
strata generate: expert cache 309 slots, 0.60 GiB of VRAM; policy is
strata serve: prompt chunk auto: 256 tokens, a 8-slot ring
strata serve: the prompt path borrows 146 CUDA0 cache slots (0.28 GiB)
strata serve: prompt 22375 tokens = 0 reused + 22375 read in 191355 ms (116.9 tok/s),
              404 generated in 21226 ms (19.0 tok/s), drafts accepted 263 of 372, 2 checkpoints
```

0.1.39 with only `--prefill 512` (start-up only, not benchmarked):

```
strata generate: expert cache 309 slots, 0.60 GiB of VRAM; policy is
strata serve: prompt chunk 512 -> 256 tokens so its buffers fit in every expert cache
strata serve: the prompt path borrows 146 CUDA0 cache slots (0.28 GiB)
```

0.1.39 with `--prefill 512 --no-prefill-borrow`:

```
strata generate: expert cache auto: 1.25 GiB free, 700 MiB reserved (+218 MiB for the draft head) -> 0 slots
strata generate: expert cache auto: the 700 MiB reserve leaves too few slots on this card
                 (a working cache needs 1): a 554 MiB reserve instead -> 1 slots
strata serve: the prompt path allocates its own buffers (too few cache slots to borrow)
strata serve: prompt 22375 tokens = 0 reused + 22375 read in 107740 ms (207.7 tok/s),
              401 generated in 18872 ms (21.2 tok/s), drafts accepted 263 of 355, 2 checkpoints
```

The log does not print the chunk for this last run. I read it as 512 because I asked for 512, there is no "512 -> 256" line, and prefill came back to the 0.1.31 level.

### Why, as far as I can tell

A 256-token chunk borrows 146 of the 309 cache slots, so a 512-token chunk would need roughly twice that, about 290. The loan has to leave 128 slots in the cache (`fits_one` in `src/program/generate.cpp`), and 290 + 128 > 309, so the chunk drops to 256. `STRATA_RING_BYTES=0` does not change this (prefill stayed at 114 tok/s).

### Workaround

Add `--prefill 512 --no-prefill-borrow` to the config's args. On this PC the expert cache then shrinks to 1 slot (the decode hit rate went from 14-22% to 0.1%), but decode did not get slower (21.2 vs 19.0 tok/s on the 22K prompt). Setup does not carry hand-added args over when it rewrites the config, so this has to be re-added after each update.

### Caveats

- One run per config. Between the two 0.1.31 runs, prefill differed by 1-4% and decode by up to 20%.
- The amount of experts held in RAM by `--resident-experts` was not the same across runs (0.1.31: none, it fell back to mmap; 0.1.39 default: 15.7 GiB; `STRATA_RING_BYTES=0`: 21.0 GiB; the workaround: 18.5 GiB), because free RAM at start differed. Prefill did not follow it (113.9-116.9 tok/s at 15.7 and 21.0 GiB).
- Not tested on a 12 GB or larger card. I expect a larger cache leaves room for a bigger borrowed chunk there.

### Possible change

When owned buffers allow a larger chunk than the borrowing plan can get, the auto choice could take the owned path, as on 0.1.31. #658 proposes something close to this (measured on an RTX 5090). Related: #765, #796.

### How it was measured

For each config the server was restarted, then one warm-up request, a 56-token code request, an ~8K prompt plus one follow-up, and a ~22K prompt plus one follow-up, all through `/v1/chat/completions` with streaming, `reasoning_effort: medium`, `temperature: 0`. The long prompts are Strata's own source files from 9259cad followed by a short question, so they are nearly the same token count on both engines. The OS file cache was not controlled.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.