Pull requests / #1739
#1739 bench: RTX 4070 12 GB / 32 GB DDR4, Flash-Next IQ2_XS at 48.23 tok/s
open · @Hcl192088 · 0 comentarios · En GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows
Descripción
## Summary Community **accepted decode throughput** benchmark for **Qwen3.8-Flash-Next IQ2_XS** on an **RTX 4070 12 GB + 32 GB DDR4-3200 + i5-12600KF** Windows 11 machine. **The benchmark uses a locally built, Adrian-derived Strata engine**, not an unmodified official upstream binary. This PR adds **benchmark documentation, measured results, and reproduction tools only**; it does **not** change the upstream engine. ## Measured results (2026-10-08) | Configuration | Valid runs | Decode tok/s, per run | Median | | --- | ---: | --- | ---: | | E4/S24 (`--adapt-every 4 --adapt-swaps 24`) | 5 | 48.23, 47.98, 49.44, 48.83, 46.92 | **48.23** | | Frozen E4/S32 reference | 5 | 45.94, 47.22, 50.04, 47.76, 49.37 | **47.76** | E4/S24 was **+0.98%** relative to the frozen reference median, **below the campaign's +1% promotion criterion**; it was not promoted. - **Actual prompt length:** 28,912 tokens; **generated:** 1,024 tokens per run. - **Configured maximum context:** 100,000 tokens (**not** a 100,000-token prompt benchmark). - **Output metric:** accepted/generated decode tokens per second reported by the engine; not speculative drafts or prompt throughput. - Model: original ISTA-DASLab GSQ-RCO **IQ2_XS**, two pinned GGUF shards. - Full hardware/software/settings, model and artifact SHA-256 identities, six environment overrides, five-run records and the exact historical command are included in `provenance.json`. ## Engine source and reproduction The **complete historical engine source** is available at [the pinned source branch](https://github.com/Hcl192088/Strata/tree/benchmark-iq2-source-9ec3806), **commit `9ec3806058cf32ab27a55e4377daf7cf0d087dec`**, based on [AdrianBM96/Strata3060](https://github.com/AdrianBM96/Strata3060). It is linked separately rather than submitted as an unrelated bulk engine diff. Included in this results PR: - [`README.md`](bench/results/2026-10-08-community-rtx4070-iq2-xs/README.md): hardware, configuration and measured results. - [`REPRODUCE.md`](bench/results/2026-10-08-community-rtx4070-iq2-xs/REPRODUCE.md): source build, pinned model/pack/MTP preparation, profile-generation methods, benchmark protocol and limitations. - [`reproduce.py`](bench/results/2026-10-08-community-rtx4070-iq2-xs/reproduce.py): local SHA-256 checks and five-run measurement harness, retaining per-run diagnostics. - [`make_local_prompt.py`](bench/results/2026-10-08-community-rtx4070-iq2-xs/make_local_prompt.py): generate a private 28,912-token input from a tester's own text. - [`provenance.json`](bench/results/2026-10-08-community-rtx4070-iq2-xs/provenance.json) and [`transcribed-runs.csv`](bench/results/2026-10-08-community-rtx4070-iq2-xs/transcribed-runs.csv): run identities and source measurements. Strata's built-in opt-in `--expert-profile-save` / `--expert-profile` and `tools/make_profile.py` allow testers to train or generate their **own** expert cache profiles. Neither the original learned profile nor the original private prompt is required to repeat the **method**. ## Limitations - The **48.23 tok/s median is an observed result on the above machine**, not a guarantee for other prompts, profiles or hardware. An independent test with a different prompt/profile is **comparable**, but not a byte-identical replay of these five runs. - Original prompt/profile bytes and full raw desktop logs have **not** been redistributed; original identities and recorded per-run measurements remain documented. - The historic binary's exact rebuild toolchain was not fully captured; compiling the pinned source can yield a different binary SHA-256. - **12 GB** denotes installed GPU VRAM; **peak inference VRAM was not measured**. - The included reproduction scripts have **not yet been independently run on another machine**. Submitted per `docs/COMMUNITY_BENCHMARKS.md`, with benchmark material separated from engine source changes.
En el sitio
Enlaces a install, modelos, releases.