Pull requests / #777
#777 Add an RTX 5090 + Ryzen 7 9700X IQ3_XXS benchmark on engine 0.1.39
closed · @steve8697 · 0 comentarios · En GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
## What this is Results-only community benchmark for the published Strata **0.1.39** engine (`6f32ec0`) on a native Windows PC. No engine changes. Folder: `bench/results/2026-10-04-community-rtx-5090-9700x/` This is a different machine from `bench/results/2026-09-30-community-rtx-5090/` (that report is Linux, Core Ultra 9 285K, 64 GB, IQ2_XS, engine 0.1.29). ## Hardware and model - RTX 5090, 32,607 MiB, power limit 402.50 W, driver 617.14. Busy samples were PCIe Gen4 x16. Clocks were not fixed. - Ryzen 7 9700X, 16 logical processors, 7 expert-pool workers. About 93.7 GiB RAM reported by the server. - Windows 11 26H2 build 26300. Published engine zip, CUDA 13.0, including the GPU vision helper. - Original Flash-Next **IQ3_XXS**, GGUF revision `ed59f92082b1e93c0e96d60a8b11aab089b52f09`. Context 262144. KV int8, `--kv-resident 32768`. Vision encoder loaded; no images were sent. ## Configuration that affects the numbers - Expert cache auto: 12,758 experts / 20.71 GiB at ready. `STRATA_PF_FUSED=1`. - Config `--spec 4` is reported by the engine as `spec=6`, `mtp_max=4`, `lookup=3`. - `"parallel": 2` reserved two slots. Requests were serial, so each stayed on the solo MTP path. - Conversation cache 8192 MiB / 4 slots. All nine timed runs reported 0 reused tokens. - An unrelated download was writing to the same NVMe during the run (8.54 GB to 12.05 GB of a 54.82 GB shard). ## Method and results Same shape as the 2026-09-30 script: one warmup, then three fresh prompts at 4096, 32768, and 128000 tokens. Temperature 0, `reasoning_effort` none, `max_tokens` 256, streaming. Prompt tok/s is fresh tokens / `prompt_ms`. Decode tok/s is `engine_generated` / `decode_ms`. | Prompt tokens | Prompt tok/s median (min–max) | Decode tok/s median (min–max) | TTFT s median (min–max) | | ---: | --- | --- | --- | | 4096 | 3775 (2617–3825) | 171.7 (130.2–173.9) | 1.11 (1.09–1.59) | | 32768 | 6367 (6221–6382) | 181.2 (157.3–183.0) | 5.22 (5.19–5.34) | | 128000 | 6080 (6037–6241) | 178.3 (174.3–193.3) | 21.24 (20.70–21.40) | The first 4K run is the slow end of that row (expert-cache hit 95.1%; later long runs sat near 99%). `tools/needle_bench.py` at 32k and 128k, depths 10/50/90: **6/6 exact**. Per-run JSON is in the folder. `COMMUNITY_BENCHMARKS.md` is unchanged so maintainers can decide whether to index this.
En el sitio
Enlaces a install, modelos, releases.