Pull requests / #1351
#1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
open · @1314521gjy · 0 comentarios · En GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descripción
Results-only report, no engine changes. **Report:** `bench/results/2026-10-04-community-rtx-4080s-iq3s/README.md` **Rebuilt on `82f46a8`:** the previous PR for this report (#780) was closed automatically when `main` was force-pushed during the history cleanup; this is the same content cherry-picked onto the new main, plus one addition (below). **Hardware/software:** RTX 4080 SUPER (32 GB, driver 616.56, PCIe Gen 4 x16), Core i7-13790F (8P+16E, AVX-2 only), 96 GB RAM, Windows 11. Engine **0.1.39**, Flash-Next **IQ3_S**, INT8 KV, vision on. **What was measured** 1. **Expert-pool workers.** Four-arm ABBA on the current configuration, 4 workers against the engine default **15** on this CPU (which is what #642's `max(1, p-1+e//2)` gives for 8P+16E). Four-arm means **95.05 vs 73.84 tok/s (+28.7% for 4)**, 8 of 8 workload-by-pair cells favour 4, and the same-configuration spread is 0.21% on the pool-4 side against 12.6% on the pool-15 side. A 0.1.36 sweep is included: the pool's own cost rises monotonically from 6.38 to 17.87 ms/round between 4 and 15 workers while cold prefill stays flat. 2. **`STRATA_PF_FUSED=1` with IQ3_S experts:** cold prefill **+8.9%** (8/8 cells positive), decode not distinguishable. 3. **Context tiers at a fixed prompt.** 524,288 costs **+0.9...+6.3%** on a fresh 43,969-token read over 262,144. 1,048,576 costs **7x** here when an explicit `--expert-cache` left the card with **0 MiB free** - on Windows an over-committed card pages instead of failing - and at 1,048,576 with a cache that fits: `auto` (11,178 slots, 217 MiB free) = 102-122 tok/s, `auto` + `--vram-reserve-mib 1500` (10,766 slots, 1,075 MiB free) = 88.9-101.4, explicit 7000 (9,148 slots, 4,130 MiB free) = 94-97, against 108-113 at 524,288. **453 slots (0.86 GiB) is the whole difference**, and the slower arm had the higher cache hit rate, so it is not misses. 4. **A third-party reproduction of the 1M cost** (gucasbrg, #781) on 2x RTX 5090 (Linux, layer split): 131,072 = 165.2 tok/s, 1,048,576 = 157.7 (**-4.4%**), YaRN free, and raising the reserve from 700 to 1500 (699/199 MiB free -> 1,197/1,071 MiB free) **did not change decode** (157.8). So the tier is cheap but not free at healthy margins, and the 7-8x here is the 0 MiB regime - the report keeps those two apart. 5. **A negative result kept as such:** the `\GPU Process Memory(*)\Shared Usage` / `Dedicated Usage` counters **did not** separate the slow arm from the fast ones on this machine (they track the pinned expert arena and pinned K/V). **Limitations:** one machine, one quantization; **no needle/recall test was run at any tier**; the first tier ladder is n=1 per arm and the cache-budget table is n=2-3 per arm on one set of three prompts; the pool sweep rows are single-arm screens on 0.1.36 while the ABBA is on 0.1.39. Attached: launch configuration, artifact hashes, per-run records (`results.json`), the engine's startup banner, and a `/metrics` sample.
En el sitio
Enlaces a install, modelos, releases.