反馈 / #1682
#1682 [Bug]: Linux file tier: the automatic read choice picks the file cache where forced unbuffered reads decode 36 % faster (v0.1.41, RAM budget, PCIe 5.0 NVMe)
open · @KarlGrier · 0 评论 · 去 GitHub 看
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindowsLinux
说明
### What happened On this Linux PC the file tier's automatic read choice (`file_tier_unbuffered`, the #1194 rule) picks reads through the file cache for the experts outside a `--resident-budget-gib` copy. Forcing `STRATA_UNBUFFERED_LOAD=1` decodes 36 % faster here on v0.1.41, and was 15 % and 48 % faster on two files on v0.1.40.2. **Stock v0.1.41, measured 2026-10-09.** Qwen3.8-Flash-Next Q8_0 (unsloth, 119.5 GiB of routed experts, more than this PC's RAM). RTX 5080 alone, `--resident-budget-gib 64`. `strata generate --stats`, greedy, 1,024 tokens after a 136-token math prompt. Model page cache dropped before each run, arms in A B B A order: | arm | the engine's choice | decode tok/s | ms per verify round | file tier: MB read per round / ms reading per round | share of those bytes not in the page cache | |---|---|---|---|---|---| | A | through the file cache (the rule) | 11.97 | 204.7 | 595 / 284 | 29 % | | B | unbuffered (`STRATA_UNBUFFERED_LOAD=1`) | 15.26 | 156.0 | 599 / 96 | – | | B | unbuffered | 15.61 | 153.7 | 618 / 95 | – | | A | through the file cache | 10.69 | 222.6 | 594 / 329 | 30 % | Mean 11.33 → 15.43 tok/s (**+36 %**). The rule's line, both A runs: ``` strata generate: the file tier reads through the file cache (84.2 GiB available, 64.0 GiB of it still to be taken by the RAM copy, 55.5 GiB of experts read from the files: the file cache can keep the ones that come back) strata generate: the file tier reads through the file cache (re-checked with the RAM copy built (64.00 GiB): 20.2 GiB available, 0.0 GiB of it still to be taken by the RAM copy, 55.5 GiB of experts read from the files: the file cache can keep the ones that come back) ``` So the cache has about 20 GiB for 55.5 GiB of experts read from the files. The rule asks it to hold 5 % of them, and 29-30 % of the bytes still came from the drive. During the buffered runs swap in use rose to 5.5-5.6 GiB (2.0 GiB before the first run); during the unbuffered runs it peaked at 3.4 GiB. **v0.1.40.2, 2026-10-07** (a local build whose changes do not touch the file tier; same rule, same method): | file | routed experts | through the file cache (rule) | `STRATA_UNBUFFERED_LOAD=1` | | |---|---|---|---|---| | Q8_0 (as above) | 119.5 GiB | 10.69 tok/s | 15.79 tok/s | +47.7 % | | a local file: UD-Q5_K_XL's routed experts with IQ3_S's other tensors | 91.6 GiB | 30.71 tok/s | 35.38 tok/s | +15.2 % | Forcing the file cache (`STRATA_UNBUFFERED_LOAD=0`) changed little or was slower there (+0.8 %, −8.3 %). This looks like the other side of #1194. There the cache won (a 31 GB PC reading from an NTFS volume, and a Tesla P100 box under 20-32 GB cgroups). Here a PCIe 5.0 NVMe drive with 64 I/O threads reads the misses faster than a 20 GiB file cache can serve them. Could the rule take the drive's speed into account, for example by timing a few miss reads both ways at start, or by preferring direct reads when the cache can hold only part of what is read (here about 20 GiB for 55.5 GiB)? Not tested: other drives, less RAM, the default `STRATA_IO_THREADS` (this config sets 64), `--mmap-experts` without a budget, Windows, and files whose experts all fit in RAM (the rule does not act there). ### Strata version, GPU, OS v0.1.41 (`fb58e0db`) source build, CUDA 13.3 · RTX 5080 16 GB (alone in this placement) · Linux 7.0 (Ubuntu 24.04-based), Ryzen 9 9950X3D, 96 GB DDR5, Samsung 9100 PRO 2 TB (PCIe 5.0 NVMe) ### Engine log Flags: `--expert-cache auto --prefill auto:16384 --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768 --pcie-frac 0 --ple-inflight 128 --vram-reserve-mib 813 --resident-budget-gib 64`; env `STRATA_IO_THREADS=64 CUDA_MODULE_LOADING=LAZY STRATA_GR_V3=1 STRATA_PF_FUSED=1 STRATA_MMVQ_IL=1 STRATA_ADAPT_LAG=2 STRATA_DISJOINT_ADAPT=1 STRATA_IO_STATS=1`. ``` strata generate: expert cache 1589 slots, 7.73 GiB of VRAM FileExpertSource: RAM budget 64.00 GiB: 13158 of the 111.80 GiB of experts the GPU cache does not hold, by profile rank; the rest are read from the files ``` Measured and drafted with Claude Code.
本站相关内容
相关页面的快捷入口。