Pull requests / #721

#721 Community benchmark: 2× RTX 3060 12 GB, Qwen3.8-Flash-Next IQ3_S (three prompt sizes, 3 runs each)

closed · @zimuhuan-code · 0 comentários · No GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Descrição

Adds a community benchmark report for Qwen3.8-Flash-Next **IQ3_S** on **2× RTX 3060 12 GB**
(Strata 0.1.37, commit `db4f91a`, built from source), following `docs/COMMUNITY_BENCHMARKS.md`.

**What is included**

- `README.md` — the report, in the guide's template (all six sections): hardware/software,
  model/configuration, method, results, correctness & limitations, plus a documented data format.
- `results.json` — every measured run, with the engine's own `timings` block verbatim
  (`prompt_n`, `cache_n`, `prompt_ms`, `prompt_per_second`, `predicted_n`, `predicted_ms`,
  `predicted_per_second`, `draft_n`).
- `samples.jsonl` — GPU/RAM sampled every 2 s during the benchmark.
- `bench.py` — the runner (standard library only).
- `engine.log` — the engine's own timing lines for the measurement session.
- `artifact-sha256.txt` — hashes for the weights, pack, tokenizer, engine binary and MTP artifacts.
- `raw/` — the unmodified API response of every measured run.
- `strata-config.json` — the launch configuration as loaded (local paths shortened).

**Summary** — median of 3 runs per size, greedy, 256-token output cap, and `cache_n = 0` on every
run (each run carried a distinct prompt prefix, so nothing was reused):

| Actual prompt tokens | Prompt tok/s (median, range) | Decode tok/s (median, range) | TTFT | Total latency |
| ---: | --- | --- | ---: | ---: |
| 3,197 | 603.0 (601.2–604.8) | 43.6 (42.2–44.0) | 5.4 s | 10.4 s |
| 24,798 | 1,156.3 (1,154.4–1,159.3) | 42.5 (41.4–43.1) | 21.8 s | 25.4 s |
| 98,599 | 1,323.3 (1,321.0–1,324.7) | 40.0 (37.1–40.6) | 75.5 s | 79.6 s |

Prompt throughput rises with prompt length here (603 → 1,323 tok/s) because the per-request fixed
cost is amortised; the "prefill 38–43 tok/s" figure from our earlier #602 note came from ~50–200
token prompts and is not comparable.

**Correctness** — `tools/needle_bench.py` at 32k and 128k × depths 10/50/90 %: **6 of 6 found**.

**Limitations** (also listed in the README): synthetic prompts (one filler paragraph repeated),
one quantisation, warm expert cache with no cold-start arm, no GPU-clock or PCIe pinning, and TTFT
measured in separate streaming runs because this version's streaming path returns no
`usage`/`timings`. A Q2_0 comparison from earlier work is quoted as **historical only** and was
explicitly **not** re-run (its weights and pack are no longer on disk).

Results-only: this PR adds a report and its data — **no engine changes**.

No site

Links install, modelos, releases.