Pull requests / #499
#499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
closed · @homeofe · 0 comentários · No GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
Results-only PR, following `docs/COMMUNITY_BENCHMARKS.md`. It is the report `docs/AMD_HIP.md` asks for: **the ready-made Windows AMD engine (`strata-windows-x64-hip.zip`, 0.1.35) running a model on a discrete card**, an RX 9070 XT 16 GB on Windows 11. Nothing was compiled; setup downloaded the zip, and the engine loaded the bundled `engine\amdhip64_7.dll` (#468). The PR changes only `bench/results/2026-10-02-community-rx9070xt-windows/` and adds one bullet to the "Community reports" list in `docs/COMMUNITY_BENCHMARKS.md`. ## Results (engine timing lines, median [range] of 3 runs) Qwen3.8-Flash-Next IQ3_S, 65,536-token context (setup's choice), KV int8, MTP on, images off, fresh server, every prompt fully read (`0 reused`): | Test | Prompt tokens | Prompt read tok/s | Decode tok/s | Drafts accepted | Decode expert cache hits | |---|---:|---:|---:|---:|---:| | 256-token answer | 72 | - | **45.2** [37.1-46.4] | 84.5% [84.1-86.3] | 84.6% [67.0-86.2] | | prompt 4K | 3,927 | **370.4** [366.8-370.9] | - | - | - | | prompt 32K | 31,059 | **601.5** [595.4-602.7] | - | - | - | - `strata-device --list-devices` / `--selftest`: device 0, gfx1201, 15.9 GiB, wave32; **selftest OK**. - VRAM: expert cache 4,983 slots (9.48 GiB), MTP 835 MiB, "462 MiB of VRAM free with everything loaded". RAM: expert arena 46.84 GiB. - The session ran without errors from setup to the last request. ## Findings - **Large pages refused on Windows:** `large pages refused for 50295996416 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages`. 1314 is `ERROR_PRIVILEGE_NOT_HELD` (no "Lock pages in memory" right, the Windows default). The engine continued with 4 KB pages; I did not test whether granting the right changes the speed. - **Short prompts read slowly:** the 72-token prompt read at 14.7-22.8 tok/s, so the first token took 3.2-4.9 s. On the same CPU/RAM with an RTX 2080 Ti on Linux (CUDA, 0.1.34) the same prompt's first token took 1.6 s. - **Setup re-downloaded a model that was already there:** the model files were copied into `Strata-data` next to a fresh clone (with their `.done` marks), but `%APPDATA%\Strata\settings.json` still had a `data_dir` from an earlier install, and setup downloaded IQ3_S into that folder instead. `--data-dir` fixed it. As designed (the remembered folder wins), but checking `Strata-data` next to the Strata folder for complete files before downloading would avoid it. ## Method `bench.py` (included, standard library only) against the OpenAI endpoint: streaming, temperature 0, `"reasoning_effort": "none"`, a unique nonce per request, one warm-up request, three runs per test. Client-side numbers agree with the engine's within 1%. Included: per-run JSON, the full engine log, the `strata-device` output, and setup's config (paths shortened to `<folder>`, no API key). Not measured: correctness checks, contexts above 32K, load time, power. ## Test machine AMD Radeon RX 9070 XT 16 GB (gfx1201, Windows driver 32.0.31041.3013), AMD Ryzen Threadripper 3960X (24 cores, AVX2), 128 GiB DDR4-3200 (8 modules, quad channel), Windows 11 Pro 10.0.26200, Strata `main` at `d9ab843` (engine 0.1.35, ready-made HIP engine with ROCm 10.2.0a20260930).
No site
Links install, modelos, releases.