Pull requests / #811
#811 Add an RTX 5090 + Ryzen 7 9700X IQ3_S benchmark on engine 0.1.39
closed · @steve8697 · 0 comentarios · En GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
## What this is Results-only community benchmark for the published Strata **0.1.39** engine (`6f32ec0`) on a native Windows PC. No engine changes. Folder: `bench/results/2026-10-05-community-rtx-5090-9700x-iq3s/` This is IQ3_S on the same PC as `bench/results/2026-10-04-community-rtx-5090-9700x/` (that report is IQ3_XXS). It is also a different machine from `bench/results/2026-09-30-community-rtx-5090/` (Linux, Core Ultra 9 285K, 64 GB, IQ2_XS, engine 0.1.29). ## Hardware and model - RTX 5090, 32,607 MiB, power limit 402.50 W, driver 617.14. Busy samples were PCIe Gen4 x16. Clocks were not fixed. - Ryzen 7 9700X, 16 logical processors, 7 expert-pool workers. About 93.7 GiB RAM reported by the server. - Windows 11 26H2 build 26300.9550. Published engine zip, CUDA 13.0, including the GPU vision helper. - Original Flash-Next **IQ3_S**, GGUF revision `ed59f92082b1e93c0e96d60a8b11aab089b52f09`. Context 262144. KV int8, `--kv-resident 32768`. Vision encoder loaded; no images were sent. ## Configuration that affects the numbers - Expert cache auto: 10,733 experts / 20.37 GiB at ready. Arena 47,962 MiB. `STRATA_PF_FUSED=1`, and the startup log printed the fused int8 prompt-expert line. Windows used 4 KB pages after large pages were refused. - Config `--spec 4` is reported by the engine as `spec=6`, `mtp_max=4`, `lookup=3`. - `"parallel": 2` reserved two slot sessions (1.90 GiB). Requests were serial, so each stayed on the solo MTP path. - Conversation cache 8192 MiB / 4 slots. All twelve timed runs reported 0 reused tokens. - The weights were already on disk. No shard download was running. ## Method and results One warmup, then three fresh prompts at 4096, 32768, 128000, and 261880 tokens. The last length fills this 262144 context after the server's 8-token margin and a 256-token answer. Temperature 0, `reasoning_effort` none, `max_tokens` 256, streaming. Prompt tok/s is fresh tokens / `prompt_ms`. Decode tok/s is `engine_generated` / `decode_ms`. | Prompt tokens | Prompt tok/s median (min–max) | Decode tok/s median (min–max) | TTFT s median (min–max) | | ---: | --- | --- | --- | | 4096 | 3072 (3056–3080) | 164.3 (116.6–178.5) | 1.36 (1.35–1.37) | | 32768 | 5963 (5925–5973) | 161.3 (156.0–179.0) | 5.55 (5.54–5.60) | | 128000 | 5803 (5794–5914) | 168.4 (153.2–178.9) | 22.25 (21.81–22.28) | | 261880 | 5535 (5522–5600) | 170.5 (166.6–173.2) | 47.67 (47.11–47.77) | The first 4K run is the slow end of that row (expert-cache hit 91.4%; later runs sat between 96.9% and 98.8%). `tools/needle_bench.py` at 32k, 128k, and 256k, depths 10/50/90: **9/9 exact**. The tool skipped `262k` because 262×1024×0.98+200 is above this server's 262144 context. Per-run JSON is in the folder. `COMMUNITY_BENCHMARKS.md` is unchanged so maintainers can decide whether to index this.
En el sitio
Enlaces a install, modelos, releases.