Issues / #1194

#1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores it

closed · @willthehuman · 4 コメント · GitHub で見る

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationLinux

本文

**Engine:** 0.1.40 (locally built, Linux, `arch=compute_89`)
**Config:** Q2_0 (`native experts v3` pack), `--mmap-experts --resident-budget-gib 18 --kv-resident 32768 --conversation-cache-mib 3072 --conversation-cache-slots 4`, context 131072, `--ple-io` default
**Machine:** i7-8700, 31 GB RAM, RTX 4070 SUPER 12 GB, NVMe (models read from an NTFS pack volume)

## Symptom

Same benchmark before and after the update — greedy, "Count from 1 to 200, one number per line", `max_tokens 260`, fresh engine process each time:

- **0.1.39:** 67.4 / 58.2 / 69.6 / 59.7 / 68.3 t/s — NVMe reads **95–190 MB/s**
- **0.1.40:** 29.4 / 31.2 / 32.7 / 33.0 / 31.9 t/s — NVMe reads **657–987 MB/s**

## What the engine says

The budget itself is granted, so this is not the clamp:

```
FileExpertSource: RAM budget 18.00 GiB: 13981 of the 25.86 GiB of experts the GPU cache does not hold, by profile rank; the rest are read from the files
FileExpertSource: mapped pinned cache complement ready: resident 18.00 GiB, pinned 18.00 GiB; page-locked and mapped
strata generate: the file tier reads unbuffered (re-checked with the RAM copy built (18.00 GiB): 6.4 GiB available, 0.0 GiB of it still to be taken by the RAM copy, 13.6 GiB of experts read from the files: the file cache cannot keep them)
```

## Where the time goes

`GET /metrics` per 260-token request — the requested volume is the same on both versions, only the cache hits differ:

- ram_blobs ~19 000–22 000, file_blobs ~7 500–8 300, **file_mb ~10–11 GB** (both versions)
- with unbuffered reads the disk sees **~6.4 GB per request**; with the file cache it saw ~1 GB (measured at the block layer, 0.1.39)

## Workaround that restored full speed

`STRATA_UNBUFFERED_LOAD=0` (systemd drop-in for our service):

```
strata generate: the file tier reads through the file cache (STRATA_UNBUFFERED_LOAD=0)
```

→ 65.4 / 55.1 / 63.5 t/s at 136–203 MB/s, i.e. 0.1.39-level speed with the 0.1.40 engine and all the new code.

## Why the heuristic looks wrong for this case

`file_cache_keeps(avail, arena_bytes, read_bytes)` compares **all** file-read bytes (13.6 GiB) against available RAM (6.4 GiB) and concludes the cache is useless. But the file tier is a **hot subset**: the same experts repeat across tokens (~30 file blobs per token here, with an 82% VRAM hit-rate), and roughly 2–3 GB of page cache was enough to serve them — 0.1.39 had no O_DIRECT tier and was 2× faster on identical hardware, config and prompt.

Possible improvements, in the project's measured style:
1. With `--resident-budget-gib` set, default the file tier to buffered (or re-evaluate after the first requests, from the observed hit pattern) — the budget already guarantees the "all experts in RAM" case is not being attempted.
2. Mention the trade-off in DETAILS.md next to #286, and note `STRATA_UNBUFFERED_LOAD=0` for the "resident budget + small page cache" case.
3. Setup could pick it (`--low-ram` answers) for machines like this one (32 GB RAM, 12 GB GPU, mapped+resident).

The docs' own #773 note ("forcing unbuffered reads there re-read 163–629 GB from the drive and halved the speed") describes the same failure mode in the no-budget case; this is the with-budget analogue, and the numbers are close (halved speed, ~6× the drive traffic).

関連リンク

インストール・モデル・リリースへの站内リンク。