Pull requests / #499

#499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S

closed · @homeofe · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

Results-only PR, following `docs/COMMUNITY_BENCHMARKS.md`. It is the report `docs/AMD_HIP.md` asks for: **the ready-made Windows AMD engine (`strata-windows-x64-hip.zip`, 0.1.35) running a model on a discrete card**, an RX 9070 XT 16 GB on Windows 11. Nothing was compiled; setup downloaded the zip, and the engine loaded the bundled `engine\amdhip64_7.dll` (#468).

The PR changes only `bench/results/2026-10-02-community-rx9070xt-windows/` and adds one bullet to the "Community reports" list in `docs/COMMUNITY_BENCHMARKS.md`.

## Results (engine timing lines, median [range] of 3 runs)

Qwen3.8-Flash-Next IQ3_S, 65,536-token context (setup's choice), KV int8, MTP on, images off, fresh server, every prompt fully read (`0 reused`):

| Test | Prompt tokens | Prompt read tok/s | Decode tok/s | Drafts accepted | Decode expert cache hits |
|---|---:|---:|---:|---:|---:|
| 256-token answer | 72 | - | **45.2** [37.1-46.4] | 84.5% [84.1-86.3] | 84.6% [67.0-86.2] |
| prompt 4K | 3,927 | **370.4** [366.8-370.9] | - | - | - |
| prompt 32K | 31,059 | **601.5** [595.4-602.7] | - | - | - |

- `strata-device --list-devices` / `--selftest`: device 0, gfx1201, 15.9 GiB, wave32; **selftest OK**.
- VRAM: expert cache 4,983 slots (9.48 GiB), MTP 835 MiB, "462 MiB of VRAM free with everything loaded". RAM: expert arena 46.84 GiB.
- The session ran without errors from setup to the last request.

## Findings

- **Large pages refused on Windows:** `large pages refused for 50295996416 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages`. 1314 is `ERROR_PRIVILEGE_NOT_HELD` (no "Lock pages in memory" right, the Windows default). The engine continued with 4 KB pages; I did not test whether granting the right changes the speed.
- **Short prompts read slowly:** the 72-token prompt read at 14.7-22.8 tok/s, so the first token took 3.2-4.9 s. On the same CPU/RAM with an RTX 2080 Ti on Linux (CUDA, 0.1.34) the same prompt's first token took 1.6 s.
- **Setup re-downloaded a model that was already there:** the model files were copied into `Strata-data` next to a fresh clone (with their `.done` marks), but `%APPDATA%\Strata\settings.json` still had a `data_dir` from an earlier install, and setup downloaded IQ3_S into that folder instead. `--data-dir` fixed it. As designed (the remembered folder wins), but checking `Strata-data` next to the Strata folder for complete files before downloading would avoid it.

## Method

`bench.py` (included, standard library only) against the OpenAI endpoint: streaming, temperature 0, `"reasoning_effort": "none"`, a unique nonce per request, one warm-up request, three runs per test. Client-side numbers agree with the engine's within 1%. Included: per-run JSON, the full engine log, the `strata-device` output, and setup's config (paths shortened to `<folder>`, no API key).

Not measured: correctness checks, contexts above 32K, load time, power.

## Test machine

AMD Radeon RX 9070 XT 16 GB (gfx1201, Windows driver 32.0.31041.3013), AMD Ryzen Threadripper 3960X (24 cores, AVX2), 128 GiB DDR4-3200 (8 modules, quad channel), Windows 11 Pro 10.0.26200, Strata `main` at `d9ab843` (engine 0.1.35, ready-made HIP engine with ROCm 10.2.0a20260930).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.