Pull requests / #618

#618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokens

closed · @HarukiOtaku · 0 commentaires · Sur GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

Community benchmark per `docs/COMMUNITY_BENCHMARKS.md`.

**Hardware:** AMD Radeon AI PRO R9700 32 GiB (gfx1201), stock clocks, compute only (the
display runs on the integrated GPU); Ryzen 9 9950X (AVX-512, 15 expert-pool workers);
Windows reports 61.5 GiB total physical memory (the DIMM kit was not recorded); NVMe on `D:`.

**Software:** Windows 11 Pro 10.0.26300, AMD driver 32.0.31036.15, prebuilt Windows HIP engine
0.1.38, repo main `99f3dbd0`.

**Model:** original Flash-Next IQ2_XS (two GGUF shards, revision `ed59f920`) plus the MTP
draft layer packed by setup; vision off, speed projection off.

**Provenance:** `PROVENANCE.txt` in the folder carries the engine `BUILD.json` verbatim,
SHA-256 for both GGUF shards and `data/expert-profile.bin`, and the OS build / CPU / RAM /
driver lines as read from the machine.

**Configuration:** `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp ...
--max-context 262144 --kv int8 --kv-resident 32768 --mmap-experts --resident-budget-gib 16`.

**Results** (three runs per length, greedy, 256-token cap, reasoning off; `benchmark.py`
reproduces them):

| Prompt | Reused | Gen | Prompt t/s (median) | Decode t/s (median) | TTFT |
| --- | --- | --- | --- | --- | --- |
| 4,198 (4,195-4,199) | 0 | 256 (236-256) | 855.8 (820.7-863.5)   | 99.8 (95.3-101.3) | 4.947 s  |
| 33,467 | 0 | 256 | 1,182.1 (1181.1-1191.2) | 97.8 (95.3-101.2) | 28.385 s |
| 131,388 | 0 | 256 (249-256) | 1,231.7 (1225.0-1233.3) | 89.9 (86.2-91.0) | 106.898 s |
| 247,990 | 0 | 256 (250-256) | 1,130.1 (1127.2-1132.0) | 87.5 (75.1-87.5) | 219.842 s |

Read the medians with a band: three reps inside one session agree to <1 % at 32K/128K/248K
and ~5 % at 4K, while the same configuration in a fresh session differs by up to 11.5 % at
32K (1,182.1 vs 1,060.6) and 14.2 % at 4K (855.8 vs 749.3), so treat a median as +/-10-15 %.
Only 4K and 32K have replicate sessions - 131,388 and 247,990 rest on one session each.

Needle checks at ~33.5K and ~131.4K, depths 10/50/90 %: 6/6 exact-string recall (one run
each, so an outcome rather than a rate - the failed first attempt scored 0/6). No stall line
appears in the log and every tabled request completed (30 requests, one machine, so this is
not a "the card is stall-free" claim). The attached log holds 71 request lines: 36 measured,
a 15-row void first attempt, 18 calibration/warm-up, and 2 others. The log's `expert tiers`
counter shows the SSD read volume scaling with prompt length (1.31 MB/token at 4K, 0.80 at
32K, 0.63 at 128K, 0.60 at 248K).

Context for the numbers: `docs/AMD_HIP.md` marks long contexts beyond 16K as not validated
on RDNA4 and the maintainer's R9700 validation is Linux at 4K/16K, so this is Windows +
gfx1201 data out to 248K. Against that page's own R9700 row (Coder IQ1_M, Ubuntu 24.04 in a
KVM guest: 4K prompt 982 tok/s, 4K decode 45.5) decode here is about twice it while 4K prompt
is lower - a different model, RAM, OS and engine, and the gap is unexplained, so neither
direction should be read as a controlled comparison.

One flag was measured and shows nothing: `STRATA_PF_FUSED=1` at 4K measured +0.8 % and at
32K +4.9 % against the slower of the two `default` sessions (against the other default
session it is -5.3 %), i.e. inside the 11.5 % session-to-session spread, so no effect is
claimed. `docs/DETAILS.md` records +12 % for IQ2_XS with that flag on NVIDIA RTX 30 and newer.
The raw legs (fused, `STRATA_PLE_BATCH=0` - the #541 stall workaround, -12.8 % here - and
default repeats) are in the folder but are not part of the results.

All numbers were taken with the unbuffered 0.1.38 file tier; #577 reports 15-40 % slower
prompts in that regime on a different (NVIDIA/CUDA) system, and no buffered leg was run here.

**Scope:** this is a community report on the ready-made Windows zip, which is what
`docs/AMD_HIP.md` asks for - after one validated Windows AMD case (an RX 9070 XT, engine
compiled on that PC, 32K, 0.1.34) it states "The ready-made zip itself has not run a model on
a discrete card yet - please report." It is not project-validated support:
`docs/AMD_HIP_PERFORMANCE.md` excludes "Windows HIP, other AMD architectures" from its claims,
`docs/AMD_HIP.md` marks long contexts beyond 16K and answer-quality benchmarks as
unvalidated, and the longest prompt here (247,990) does not fill the configured
262,144-token window.

**Files:** `bench/results/2026-10-03-community-r9700-windows/` - `README.md` (the report),
`METHOD-AND-DISCLOSURES.md` (full method detail, label/session caveats and the complete
disclosure list), `PROVENANCE.txt` (engine build, model-artefact SHA-256, OS/CPU/RAM/driver as
read from the machine), `SHA256SUMS.txt` (checksums of the packet's own files), launch config,
harness, per-run JSON, engine log of the sessions.

Sur le site

Liens install, modèles, releases.