Pull requests / #777

#777 Add an RTX 5090 + Ryzen 7 9700X IQ3_XXS benchmark on engine 0.1.39

closed · @steve8697 · 0 Kommentare · Auf GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

## What this is

Results-only community benchmark for the published Strata **0.1.39** engine (`6f32ec0`) on a native Windows PC. No engine changes.

Folder: `bench/results/2026-10-04-community-rtx-5090-9700x/`

This is a different machine from `bench/results/2026-09-30-community-rtx-5090/` (that report is Linux, Core Ultra 9 285K, 64 GB, IQ2_XS, engine 0.1.29).

## Hardware and model

- RTX 5090, 32,607 MiB, power limit 402.50 W, driver 617.14. Busy samples were PCIe Gen4 x16. Clocks were not fixed.
- Ryzen 7 9700X, 16 logical processors, 7 expert-pool workers. About 93.7 GiB RAM reported by the server.
- Windows 11 26H2 build 26300. Published engine zip, CUDA 13.0, including the GPU vision helper.
- Original Flash-Next **IQ3_XXS**, GGUF revision `ed59f92082b1e93c0e96d60a8b11aab089b52f09`. Context 262144. KV int8, `--kv-resident 32768`. Vision encoder loaded; no images were sent.

## Configuration that affects the numbers

- Expert cache auto: 12,758 experts / 20.71 GiB at ready. `STRATA_PF_FUSED=1`.
- Config `--spec 4` is reported by the engine as `spec=6`, `mtp_max=4`, `lookup=3`.
- `"parallel": 2` reserved two slots. Requests were serial, so each stayed on the solo MTP path.
- Conversation cache 8192 MiB / 4 slots. All nine timed runs reported 0 reused tokens.
- An unrelated download was writing to the same NVMe during the run (8.54 GB to 12.05 GB of a 54.82 GB shard).

## Method and results

Same shape as the 2026-09-30 script: one warmup, then three fresh prompts at 4096, 32768, and 128000 tokens. Temperature 0, `reasoning_effort` none, `max_tokens` 256, streaming. Prompt tok/s is fresh tokens / `prompt_ms`. Decode tok/s is `engine_generated` / `decode_ms`.

| Prompt tokens | Prompt tok/s median (min–max) | Decode tok/s median (min–max) | TTFT s median (min–max) |
| ---: | --- | --- | --- |
| 4096 | 3775 (2617–3825) | 171.7 (130.2–173.9) | 1.11 (1.09–1.59) |
| 32768 | 6367 (6221–6382) | 181.2 (157.3–183.0) | 5.22 (5.19–5.34) |
| 128000 | 6080 (6037–6241) | 178.3 (174.3–193.3) | 21.24 (20.70–21.40) |

The first 4K run is the slow end of that row (expert-cache hit 95.1%; later long runs sat near 99%).

`tools/needle_bench.py` at 32k and 128k, depths 10/50/90: **6/6 exact**. Per-run JSON is in the folder. `COMMUNITY_BENCHMARKS.md` is unchanged so maintainers can decide whether to index this.

Mehr auf der Site

Links zu Install, Modellen, Releases.