Issues / #775

#775 Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile

closed · @JiuYue0820 · 2 comments · View on GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Description

Sharing a measured result and a writeup: **docs/TUNING-256K-16GB.md** on my fork ->
https://github.com/JiuYue0820/Strata/blob/docs-256k-tuning/docs/TUNING-256K-16GB.md

**PC:** RTX 5070 Ti 16 GB, i7-14700KF (8P+12E), 96 GB DDR5-5600, Windows, engine 0.1.38, Qwen3.8-Flash-Next IQ3_S (GSQ-RCO), vision encoder on GPU, context 262,144, KV int8 + streaming (`--kv-resident 32768`).

**Changes:** `START-HERE.bat --calibrate` (kept `--pool-workers 13`, `--spec-min-p 0.70`, `--pcie-frac 0.00`), `--spec 6`, and the learned expert profile (`--expert-profile-save`).

**Benchmark:** one 257,630-token prompt (this repo's own source/docs) through the HTTP API, greedy, `reasoning_effort: none`. "Warm" = prefix reused, which is how agent sessions actually run. Single runs, so the usual ±20% noise applies.

| Configuration | Cold | Warm |
| --- | ---: | ---: |
| Stock (`--spec 4`, no calibration, shipped profile) | 17.5 tok/s | 17.2 tok/s |
| `--spec 6` + calibrated | 43.0 tok/s | 53.5 tok/s |
| + learned expert profile | 45.5 tok/s | 47.0 tok/s |

Prompt reading went from 1,468 to 1,744 tok/s.

The calibration internals on this machine, in case they are useful data points:
- `--pool-workers`: 19 → 36.9 tok/s, **13 → 54.4**, 10 → 41.3 (E-cores stall every verify window; the defaults were measured on a 6-core Ryzen with no E-cores)
- `--spec-min-p`: 0.3 → 36.6, 0.5 → 41.7, **0.7 → 44.6**
- `--pcie-frac`: **0.0 → 42.1**, 0.2 → 37.9, 0.35 → 27.9, 0.55 → 23.9, 0.75 → 20.7 (PCIe 5 x16 + fast CPU: computing a missed expert beats copying it)

Also fills in the IQ3_S 262K row that DETAILS.md leaves unmeasured: on 96 GB of RAM it runs with room to spare, 38.8-53.5 tok/s across all post-tuning runs.

One documentation gap this exposed: a `--setup` re-run rewrites the config and silently drops a manually added `--expert-profile`/`--expert-profile-save` pair (the three calibrated settings survive via the settings file). Worth either preserving or printing a note.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.