Pull requests / #440

#440 bench: community report, RTX 5090 + Ryzen 9 9950X3D, IQ3_S at the full 262K window; prefill auto:32768 A/B; conversation-cache 32 GiB demo

closed · @Ambolio · 0 comentários · No GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Descrição

Community benchmark report under `bench/results/2026-10-02-community-rtx-5090-iq3s/`, results only (no engine changes), in the `docs/COMMUNITY_BENCHMARKS.md` format.

**Hardware:** RTX 5090 32 GB (Windows 11, WDDM, driver 617.14). The 4K desktop sits on a second GPU, so the 5090 carries no graphics clients — relevant to #279: `--vram-reserve-mib 100` shows no stalls here. Ryzen 9 9950X3D (AVX-512), 96 GB RAM, NVMe.

**Software:** engine 0.1.33 release binary (BUILD.json: CUDA 13.0, sm_75/86/89/120+PTX), serve wrapper 0.1.30.

**Model:** Qwen3.8-Flash-Next GSQ-RCO **IQ3_S** (2 GGUF shards + PLE), MTP `rt`, shipped expert profile, vision on CPU (16 threads).

**Configuration:** `--expert-cache auto --prefill auto:32768 --spec 4 --spec-min-p 0.70 --mtp rt --max-context 262144 --kv int8 --kv-resident 65536 --vram-reserve-mib 100 --pcie-frac 0.35 --conversation-cache-mib 32768 --conversation-cache-slots 8`, `fit_max_tokens` on. Expert cache 12,439 slots / 23.61 GiB. Full config in `config.json`.

**Findings:**
1. Speed sweep (community method: fresh prompt per run — `cache_n = 0` in every measured run — greedy, 256 out, 1 warm-up + 3): prefill **999 / 2,439 / 5,045 / 6,004 tok/s** at 1K/4K/32K/128K; decode **144 / 149 / 142 / 125**; drafts accepted 75-76%.
2. **A/B of `--prefill auto:32768`** (the `auto:32768` form from the help; PR #282's idea): 32K 4,168 -> **5,045 (+21%)**, 128K 4,450 -> **6,004 (+35%)**. Larger than #282's +15% at IQ2_XS — consistent with their note that bigger experts gain more. Decode unchanged.
3. **`needle_bench.py` 9 of 9 found** at 32k/128k/256k x depths 10/50/90 — the 256k rows are **261,669-261,670 prompt tokens**, the edge of the 262,144 window, on one 5090.
4. **Conversation-cache demo** (`conversation_cache_demo.py`): the engine default is off (`--conversation-cache-mib 0`, 4 slots). With 32 GiB / 8 slots, follow-ups on two interleaved ~65k-token conversations cost **0.6-1.2 s instead of 11-13 s (10-20x)**; the engine restores 66,251 tokens in 82 ms and a live 160,978-token agent session in 186 ms; after an LRU eviction the partial path still reused 30,560 tokens and read 25,500 fresh at 5,002 tok/s. Parked bytes ~ 30 KB/token; 8 conversations (largest 162k tokens, 3.85 GB) sat in ~20.8 GB.
5. Observation: with `auto:32768` the engine warns "0 MiB of VRAM free with everything loaded — add `--vram-reserve-mib 612`". The prompt-path loan (7,736 slots / 14.64 GiB) peaks during prompt processing; idle free is 485 MiB and no stalls appeared across the full sweep. Is the warning conservative for this shape, or should the loan cap account for the reserve?

**Limitations:** one machine, single stream, greedy, prompts built from repo text; TTFT measured client-side on a FIFO server (includes queue time); the baseline per-run JSON was overwritten by the A/B run (its per-run lines are in `bench.out`); no quality suite beyond the needle checks.

Measurements and write-up put together with Claude Code on my machine and checked by me.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.