贡献 / #1739

#1739 bench: RTX 4070 12 GB / 32 GB DDR4, Flash-Next IQ2_XS at 48.23 tok/s

open · @Hcl192088 · 0 评论 · 去 GitHub 看

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows

说明

## Summary

Community **accepted decode throughput** benchmark for **Qwen3.8-Flash-Next IQ2_XS** on an **RTX 4070 12 GB + 32 GB DDR4-3200 + i5-12600KF** Windows 11 machine.

**The benchmark uses a locally built, Adrian-derived Strata engine**, not an unmodified official upstream binary. This PR adds **benchmark documentation, measured results, and reproduction tools only**; it does **not** change the upstream engine.

## Measured results (2026-10-08)

| Configuration | Valid runs | Decode tok/s, per run | Median |
| --- | ---: | --- | ---: |
| E4/S24 (`--adapt-every 4 --adapt-swaps 24`) | 5 | 48.23, 47.98, 49.44, 48.83, 46.92 | **48.23** |
| Frozen E4/S32 reference | 5 | 45.94, 47.22, 50.04, 47.76, 49.37 | **47.76** |

E4/S24 was **+0.98%** relative to the frozen reference median, **below the campaign's +1% promotion criterion**; it was not promoted.

- **Actual prompt length:** 28,912 tokens; **generated:** 1,024 tokens per run.
- **Configured maximum context:** 100,000 tokens (**not** a 100,000-token prompt benchmark).
- **Output metric:** accepted/generated decode tokens per second reported by the engine; not speculative drafts or prompt throughput.
- Model: original ISTA-DASLab GSQ-RCO **IQ2_XS**, two pinned GGUF shards.
- Full hardware/software/settings, model and artifact SHA-256 identities, six environment overrides, five-run records and the exact historical command are included in `provenance.json`.

## Engine source and reproduction

The **complete historical engine source** is available at [the pinned source branch](https://github.com/Hcl192088/Strata/tree/benchmark-iq2-source-9ec3806), **commit `9ec3806058cf32ab27a55e4377daf7cf0d087dec`**, based on [AdrianBM96/Strata3060](https://github.com/AdrianBM96/Strata3060). It is linked separately rather than submitted as an unrelated bulk engine diff.

Included in this results PR:
- [`README.md`](bench/results/2026-10-08-community-rtx4070-iq2-xs/README.md): hardware, configuration and measured results.
- [`REPRODUCE.md`](bench/results/2026-10-08-community-rtx4070-iq2-xs/REPRODUCE.md): source build, pinned model/pack/MTP preparation, profile-generation methods, benchmark protocol and limitations.
- [`reproduce.py`](bench/results/2026-10-08-community-rtx4070-iq2-xs/reproduce.py): local SHA-256 checks and five-run measurement harness, retaining per-run diagnostics.
- [`make_local_prompt.py`](bench/results/2026-10-08-community-rtx4070-iq2-xs/make_local_prompt.py): generate a private 28,912-token input from a tester's own text.
- [`provenance.json`](bench/results/2026-10-08-community-rtx4070-iq2-xs/provenance.json) and [`transcribed-runs.csv`](bench/results/2026-10-08-community-rtx4070-iq2-xs/transcribed-runs.csv): run identities and source measurements.

Strata's built-in opt-in `--expert-profile-save` / `--expert-profile` and `tools/make_profile.py` allow testers to train or generate their **own** expert cache profiles. Neither the original learned profile nor the original private prompt is required to repeat the **method**.

## Limitations

- The **48.23 tok/s median is an observed result on the above machine**, not a guarantee for other prompts, profiles or hardware. An independent test with a different prompt/profile is **comparable**, but not a byte-identical replay of these five runs.
- Original prompt/profile bytes and full raw desktop logs have **not** been redistributed; original identities and recorded per-run measurements remain documented.
- The historic binary's exact rebuild toolchain was not fully captured; compiling the pinned source can yield a different binary SHA-256.
- **12 GB** denotes installed GPU VRAM; **peak inference VRAM was not measured**.
- The included reproduction scripts have **not yet been independently run on another machine**.

Submitted per `docs/COMMUNITY_BENCHMARKS.md`, with benchmark material separated from engine source changes.

本站相关内容

相关页面的快捷入口。