Pull requests / #1377

#1377 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.2

open · @Dmitry-B · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux

描述


Two community reports from the same PC and the same server configuration as the earlier
[2026-10-03](https://github.com/Niko1221/Strata/tree/main/bench/results/2026-10-03-community-rtx4090-iq3xxs-200k-code)
and [2026-10-04](https://github.com/Niko1221/Strata/tree/main/bench/results/2026-10-04-community-rtx4090-iq3xxs-200k-code)
reports: engine **0.1.40.2**, measured 2026-10-07.

- `bench/results/2026-10-07-community-rtx4090-iq3xxs-200k-code/` — code-explanation prompts
- `bench/results/2026-10-07-community-rtx4090-iq3xxs-200k-ru/` — Russian prose, same sweep

## Hardware, model, configuration

RTX 4090 24 GB (23028 MiB), Ryzen 9 7950X, 46464 MiB RAM, NVMe (ADATA LEGEND 960), Ubuntu 26.04.1,
kernel 7.0.0-38, driver 610.57.04, CUDA 13.4. Model `qwen3.8-flash-next-iq3_xxs`, context 204800,
INT8 KV, `--mmap-experts`, expert cache auto, MTP draft pack, `--spec 4`, `draft_vocab=cyrillic`,
experimental speed projection (control vector, layers 4-44), `--vram-reserve-mib 989`, `--pool-workers 10`.
Engine **compiled from source** for `sm_89` with CUDA 13.4 — the release publishes no Linux engine
archive for this tag, so this is the only path on Linux; both compared versions were built the same way.

Method is unchanged from the previous reports: full warm-up sweep discarded, then 4K / 32K / 128K and a
generation-only case, 3 runs each, random marker in every prompt (no prefix reuse), throughput from the
engine's timing fields, TTFT over streaming, `tools/needle_bench.py` at 32K and 128K, depths 10/50/90.
The engine's auto prompt chunk is `8192 tokens, a 96-slot ring` in both versions, so the numbers are
comparable with the 0.1.38 / 0.1.39 reports from this PC.

## What the reports contain

Each folder has the 0.1.40.2 run (`runs.json`), the **0.1.40 baseline** of the same sweep
(`runs-0.1.40-baseline.json`, measured 2026-10-06) and an **identical-binary control**
(`runs-0.1.40.1-control.json` — 0.1.40.1 changed only the Python server, `engine/strata` was not rebuilt).
The control run is there to make the noise band explicit instead of asking reviewers to guess it:
prompt -1.1…-2.0%, decode -3.1…-5.6% (code variant) and prompt -0.7…+0.8%, decode -0.1…+8.6% (Russian variant).

## Results, in short

- **Prompt throughput is faster: +5.7…10.2% (code) and +3.6…16.4% (Russian) at every length**, TTFT lower by
  0.03-2.35 s. The prompt column is tight in both versions and moved by at most 2% (0.8% in the Russian
  variant) in the identical-binary control, so the direction is solid. The configuration uses
  `--mmap-experts`, which fits the stager wait change (#1057).
- **Decode: no measurable change.** Every case stays inside the control band. The F4 verify windows
  (`STRATA_MMVQ_IL`, on by default for RTX 30+, bit-identical) did not move decode on this card.
- Recall 6/6 at 32K and 128K, unchanged.
- The Russian variant is published as a second folder on purpose: prompt throughput is essentially
  script-independent (3351-3383 vs 3310-3369 tok/s at 32K), decode is not (89.9-108.1 vs 133.5-141.7 tok/s),
  and the gap tracks draft acceptance (43-62% accepted vs 72-79%). Cyrillic text also tokenizes ~2% denser
  than the code text at the same nominal length.

## Limitations stated in the reports

Synthetic prompts, greedy decoding, early stops on repetitive text (actual generated lengths are in the
tables), the OS not isolated, and this server also serves an interactive agent session — no request from it
was issued during the measured runs, and the one that landed in the discarded warm-up is described there.
Opt-ins from this release (`STRATA_PREFILL_CPU_SHARE`, `STRATA_IO_PREFETCH`, `pin=N`, `STRATA_SPEC_GUMBEL`,
`STRATA_MMVQ_IL=0`) were not tested and are not part of these numbers. The measurement script is a local
script, not included in the PR; available on request.

No engine changes in this PR — measurements only.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。