Pull requests / #1216

#1216 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens (reland)

closed · @Zauberio · 0 Kommentare · Auf GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsWindows

Beschreibung

Reland of the community report that `main`'s force-push auto-closed (#832) — same content, rebased onto the current `main` per your note, with one appendix row corrected.

## The report

`bench/results/2026-10-04-community-rtx-5070ti-windows/` — a single RTX 5070 Ti (16 GB) + 64 GB RAM on native Windows 11, IQ3_S native pack, 262K context.

- Battery at 5K / 41K / 131K / 164K prompt tokens: numbers are the engine's own per-request log lines, not client walls.
- Model files hash-verified against the HF revision's published LFS OIDs (`model-provenance.json`); engine digest verified against the release asset at install.
- Harness included (`battery.ps1`).
- The appendix condenses this box's A/B history since 0.1.35 (version ladder, kept levers, measured losers with numbers).

## What changed since #832: the `STRATA_GR_DOWN_MAX4` row

The appendix previously credited `STRATA_GR_DOWN_MAX4=1` with +2.8% / +3.5% decode. **That claim is withdrawn** — it does not reproduce, and we believe we know why.

Our battery prefixes each run with a unique random marker to defeat the prompt cache. That also makes the two arms generate *different text*, so they get different draft-acceptance and different decode speed — a variable the historical ABBA averaged over but did not cancel.

Re-tested with a paired design (flag alternating within a pair, prompt **identical in both arms**, drifting only between pairs; decode from the engine log):

| quant | 4.5K | 22.6K | 99K | pairs |
|---|---|---|---|---|
| IQ3_S | 1.026 | 1.008 | 1.021 | 10 |
| Q2_0 | 1.004 | 0.997 | 1.000 | 8 |

Medians sit at 1.0 and the per-pair ratios scatter ±10% and flip sign. A positive control with the same harness on `STRATA_PF_FUSED` (0 vs 1) read prefill **1.023, 4/4 pairs positive** — so the method does resolve a real ~2.3% effect; this is an absence, not a resolution limit. Resolution is ~2–3%; a sub-1% real effect is not excluded.

This agrees with the maintainer's own paired measurement (0.97 / 0.97 / 0.99 on Q2_0) that reverted the sm_120 default. The row now reads "no reproducible effect" and the lever is no longer cited as evidence for turning it on by default. Everything else in the report is unchanged.

## Q2_0

To run this we built a Q2_0 pack for the 16 GB card (`C:\Strata-data\packs\q2_0`, from the ISTA-DASLab GGUF; both shards' sha256 verified against the HF LFS oids — shard 2 is byte-identical across quants). Raw per-pair CSV available on request.

Mehr auf der Site

Links zu Install, Modellen, Releases.