Pull requests / #1467

#1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3

open · @Dmitry-B · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

描述


Follow-up of [#1377](https://github.com/Niko1221/Strata/pull/1377) / [2026-10-07](https://github.com/Niko1221/Strata/tree/main/bench/results/2026-10-07-community-rtx4090-iq3xxs-200k-code): the same PC, the same model, the same server config and the same method, now on engine **0.1.40.3**, measured 2026-10-08.

- `bench/results/2026-10-08-community-rtx4090-iq3xxs-200k-code/` - code-explanation prompts
- `bench/results/2026-10-08-community-rtx4090-iq3xxs-200k-ru/` - Russian prose, same sweep

## Hardware, model, configuration

RTX 4090 24 GB (23028 MiB), Ryzen 9 7950X, 46464 MiB RAM, NVMe (ADATA LEGEND 960), Ubuntu 26.04.1, kernel 7.0.0-38, driver 610.57.04, CUDA 13.4. Model `qwen3.8-flash-next-iq3_xxs`, context 204800, INT8 KV, `--mmap-experts`, expert cache auto, MTP draft pack, `--spec 4`, `draft_vocab=cyrillic`, experimental speed projection (control vector, layers 4-44), `--vram-reserve-mib 989`, `--pool-workers 10`. Engine **compiled from source** for `sm_89` with CUDA 13.4 - the release publishes Windows archives only, so this is the only path on Linux; 0.1.40, 0.1.40.1, 0.1.40.2 and 0.1.40.3 in these reports were all built the same way on this machine.

Unchanged from the previous reports: the launch command (compared line by line in the server log), the server config files, the learned expert profile, and the engine's auto prompt chunk - `8192 tokens, a 96-slot ring` in 0.1.38 through 0.1.40.3 - so the numbers are comparable with the 0.1.38 / 0.1.39 / 0.1.40 / 0.1.40.2 reports from this PC.

Method is unchanged: full warm-up sweep discarded (its first 4K read ran at 1660.6 tok/s with a 2.416 s TTFT against 2146-2162 tok/s and 1.86-1.87 s in the measured pass - the cold expert cache, not the version), then 4K / 32K / 128K and a generation-only case, 3 runs each, random marker in every prompt (no prefix reuse), throughput from the engine's timing fields, TTFT over streaming, `tools/needle_bench.py` at 32K and 128K, depths 10/50/90.

## What each folder contains

`runs.json` (0.1.40.3) plus three reference runs, so the comparison is self-contained:

- `runs-0.1.40.2-previous.json` - the previous version, measured 2026-10-07 with this method;
- `runs-0.1.40-baseline.json` - the 0.1.40 baseline, measured 2026-10-06;
- `runs-0.1.40.1-control.json` - a full repeat of the sweep on **the identical engine binary** (0.1.40.1 changed only the Python server). This is the noise band of the method on this machine: prompt -6.6…+0.8%, decode -5.6…+8.6%.

## Results, in short

- **0.1.40.3 is not measurable on this machine.** Prompt throughput moved by -2.8…-0.6% (code) and -0.2…+3.6% (Russian) against a same-binary band of -6.6…+0.8%; decode moved by -5.0…+3.5% and -5.6…+3.2% against a band of -5.6…+8.6%; TTFT by 0.03-0.24 s. The ranges in both versions are tight (3278-3314 and 3335-3341 tok/s at 131072 tokens), so this is a negative result rather than an inconclusive one.
- **That matches what the release changes.** For a Linux/NVIDIA build the diff touches: an `n_expert == 512 && K == 10` guard on the drafter's per-token router call (#1357) - this model has `expert_count = 512` and `expert_used_count = 10` in its GGUF metadata, so the guarded native path is the one already taken; HIP/Windows-only diagnostics in `generate.cpp`; a kernel self-test in `native_multi_parity.cpp`; and the commit that made `STRATA_MMVQ_IL` opt-in is reverted in the same release, so the interleaved q8_1 projections stay on by default exactly as in 0.1.40.2. Everything else in 0.1.40.3 is Intel Arc setup, Windows AMD packaging, Docker, the web app and the tokenizer.
- **The 0.1.40 -> 0.1.40.2 prompt gain is intact.** Against the 0.1.40 baseline this build reads +2.7…+8.5% (code) and +6.4…+16.1% (Russian) with TTFT 0.02-2.5 s lower. Trend at 131072 tokens on this PC: 2962 -> 3085 -> 3119 -> 3313 -> 3292 (code) and 2984 -> 3126 -> 3136 -> 3317 -> 3336 (Russian) tok/s across 0.1.38 / 0.1.39 / 0.1.40 / 0.1.40.2 / 0.1.40.3.
- **One correction to #1377.** That report attributed the 0.1.40.2 prompt gain to the stager wait change (#1057). The 0.1.40.3 docs measure the opposite - sleeping stager waits read prompts 5-6% slower on a Ryzen 9 7940HS + RTX 4070 laptop (while whole-machine CPU use falls from 77-90% to 21-25%), and spinning is 1.2% faster on a desktop RTX 5070 - so #1057 is a CPU-load feature, not a speed feature. That attribution in the 0.1.40.2 report should be read as withdrawn; its measured numbers are unchanged. The cause of the 0.1.40.2 prompt gain is not identified by these measurements - only that it is uniform across lengths.
- Recall 6/6 at 32K and 128K, unchanged.
- The Russian variant is published as a second folder on purpose: prompt throughput is essentially script-independent (3363-3383 vs 3316-3351 tok/s at 32K), decode is not (95.4-100.9 vs 136.2-138.9 tok/s), and the gap tracks draft acceptance (47-80% vs 74-79%). Cyrillic also tokenizes ~2% denser than the code text at the same nominal length, so cross-script comparisons have to use actual token counts.

## Limitations stated in the reports

Synthetic prompts, greedy decoding, early stops on repetitive text (actual generated lengths are in the tables - and they are why the decode column is not comparable across versions when the lengths differ), the OS not isolated, and this server also serves an interactive agent session - no request from it was issued during the measured runs, and the discarded warm-up is described in each report. Opt-ins from this release (`STRATA_PREFILL_CPU_SHARE`, `STRATA_IO_PREFETCH`, `pin=N`, `STRATA_SPEC_GUMBEL`, `STRATA_MMVQ_IL=0`, `STRATA_STAGER_SLEEP=0`) were not tested and are not part of these numbers. The measurement script is a local script, not included in the PR; available on request.

No engine changes in this PR - measurements only.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。