Pull requests / #780

#780 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep

closed · @1314521gjy · 0 コメント · GitHub で見る

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

本文

Results-only report, no engine changes.

**Report:** `bench/results/2026-10-04-community-rtx-4080s-iq3s/README.md`

**Hardware/software:** RTX 4080 SUPER (32 GB, driver 616.56, PCIe Gen 4 x16), Core i7-13790F (8P+16E, AVX-2 only), 96 GB RAM, Windows 11. Engine **0.1.39**, Flash-Next **IQ3_S**, INT8 KV, vision on.

**What was measured**

1. **Expert-pool workers.** Four-arm ABBA on the current configuration, 4 workers against the engine's default **15** on this CPU (which is also what #642's `max(1, p-1+e//2)` gives for 8P+16E). Four-arm means **95.05 vs 73.84 tok/s (+28.7% for 4)**, 8 of 8 workload-by-pair cells favour 4, and the same-configuration spread is 0.21% on the pool-4 side against 12.6% on the pool-15 side (two of its four workload medians collapsed to 57-58 tok/s). A 0.1.36 sweep is included for the shape of the effect: the pool's own cost rises monotonically from 6.38 to 17.87 ms/round between 4 and 15 workers while cold prefill stays flat.
2. **`STRATA_PF_FUSED=1` with IQ3_S experts:** cold prefill **+8.9%** (8/8 cells positive, pairs +11.1%/+6.9%), decode not distinguishable. An earlier measurement said +12.8%; the re-measurement revises the size.
3. **Context tiers at a fixed prompt.** 524,288 costs **+0.9...+6.3%** on a fresh 43,969-token read over 262,144. 1,048,576 costs almost nothing **if the expert cache fits** - and this is the part that changed our conclusion: with an over-sized explicit `--expert-cache`, the same tier ran **7x slower**, because on Windows an over-committed card pages to system memory instead of failing. At 1,048,576: 11,631 slots with **0 MiB free = 13.7 tok/s**, `auto` (11,178 slots, 217 MiB free) = 102-122 tok/s, `auto` + `--vram-reserve-mib 1500` (10,766 slots, 1,075 MiB free) = 88.9-101.4 tok/s, 9,148 slots (4,130 MiB free) = 94-97 tok/s. 453 slots (0.86 GiB) is the whole difference, and the slower arm had the **higher** cache hit rate, so it is not misses. The mechanism was suggested by the maintainer in #781; the profiler tables for the three tiers and the correction to our earlier TLB/locality guess are in the report.
4. **A negative result we kept:** the `\GPU Process Memory(*)\Shared Usage` / `Dedicated Usage` counters **did not** separate the slow arm from the fast ones on this machine - they track the pinned expert arena (50.3 GB) and the pinned K/V (6.19 -> 12.38 GiB), so the fix is documented as the `--vram-reserve-mib` lever and the startup line's free figure, not that counter. Also in PR #799 as documentation.
5. The engine reports only **226 MiB of VRAM free** in the original 524,288 configuration and suggests `--vram-reserve-mib 986`; the 50.3 GB expert arena is on 4 KB pages (large pages refused).

**Limitations:** one machine, one quantization; **no needle/recall test was run at any tier**, so no recall claim is made; the first tier ladder is n=1 per arm and the cache-budget table is n=2-3 per arm on one set of three prompts; the sweep rows are single-arm screens on 0.1.36 while the ABBA is on 0.1.39.

Attached: launch configuration, artifact hashes, per-run records (`results.json`), the engine's startup banner, and a `/metrics` sample.

関連リンク

インストール・モデル・リリースへの站内リンク。