Issues / #602
#602 Measured numbers: Qwen3.8-Flash-Next on 2× RTX 3060 (Q2_0 vs IQ3_S), plus a local reproduction of #75
closed · @zimuhuan-code · 2 comentários · No GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationLinux
Descrição
Offered in the spirit of `docs/COMMUNITY_BENCHMARKS.md`. **Honest scope:** Q2_0 = **a single recorded run** (1015 tok / 22.6 s weighted mean; we ran the mix twice and report the run that matches the table); IQ3_S = **three runs, median + range**. Happy to re-run Q2_0 three times, or to emit your per-run JSON schema — tell us which fields you want.
**Environment** (structured after `docs/COMMUNITY_BENCHMARKS.md`'s checklist)
*Hardware*
- 2× NVIDIA RTX 3060 12 GB (24 GB total VRAM), driver 595.84; **power limit 170 W** (max 187 W)
- **PCIe: Gen1 ×16 when idle → Gen3 ×16 under load** (capability Gen3 ×16; measured both states — idle shows Gen1, i.e. NVIDIA's link power-saving, and it reports Gen3 during generation). So the numbers below ran on a **Gen3 ×16** link; we did not pin the link state.
- CPU Intel Xeon E5-2678 v3 @ 2.50 GHz — **2 sockets × 12 cores = 24 cores / 48 threads** (**no AVX-512**)
- RAM **128 GiB**, populated as **8 × 16 GiB DDR3-1867** — `lshw -C memory` reports `DIMM DDR3 Synchronous 1867 MHz (0.5 ns)`, `size: 16GiB` per populated slot (128 GiB ÷ 16 GiB = 8 modules, matching the 8 populated slots).
⚠️ **Platform note that matters for comparison:** this is a **modified ("magic") X99 board (Huanan Gold) running DDR3 on an LGA2011-3 Xeon E5 v3**. The stock Haswell-EP memory controller is DDR4-only — DDR3 support here is a *board-level* feature, not a CPU one. Memory bandwidth therefore need not match a stock DDR4 X99 platform, which is worth keeping in mind when comparing the prefill numbers below.
- Storage: models on **NVMe SSD** (`nvme1n1`, 1.9 TB); the machine also has 2× HDD, idle for these runs
- Other workloads: none — `ftllm` (which shares these two GPUs) is stopped while Strata runs
*Software*
- OS Ubuntu 26.04, glibc **2.43** (`ldd --version` → 2.43-2ubuntu2.3)
- **Strata v0.1.37, commit `db4f91a`** (engine 0.1.37), **built from source** (no Linux release binary existed: `releases/latest` 404)
- CUDA **12.8** (12.9 fails to build here — see #601); **changed build options**: `CXX=/usr/bin/g++-14 CUDAHOSTCXX=/usr/bin/g++-14` plus a local `setup.py` patch to pick the toolkit via `STRATA_NVCC`
- llama.cpp pinned by setup at `3cf03257f219afbe7334045ff7c6a06ac68c627d`; Python 3.14.4 (setup's venv)
- gcc/g++ 13/14/15 installed, default `g++` = 15
*Model*
- Repo **`ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`**, revision **`ed59f92082b1e93c0e96d60a8b11aab089b52f09`** (pinned in `setup.py`)
- Quantisations **Q2_0** and **IQ3_S**; files `Qwen3.8-Flash-Next-GSQ-RCO-{Q2_0,IQ3_S}-0000{1,2}-of-00002.gguf`
(shard 2 is identical between the two quantisations → hard-linked on disk, so both together cost ≈150 GB, not 66 + 83.6)
- Vision encoder **off**; pack / expert profile / MTP draft all produced by `./setup.sh` (no hand edits); experimental speed projection **off**
*Settings* — exact launch arguments (IQ3_S; Q2_0 differs only in pack/model paths)
```
--pack …/packs/iq3_s --native …/IQ3_S/…-00001-of-00002.gguf --ple-gguf …/IQ3_S/…-00002-of-00002.gguf
--expert-profile …/data/expert-profile.bin --expert-cache auto --prefill auto
--spec 4 --spec-min-p 0.5 --mtp …/mtp/rt --max-context 524288 --rope-scaling yarn --rope-scale 2
--kv int8 --kv-resident 32768
```
- context **524288**; **yarn rope scaling ×2** (past the trained 262144); KV **int8**, streamed to RAM above 64K (**2.58 GiB pinned** at 512K, `--kv-resident 32768` cells kept in VRAM)
- expert cache **auto** → 5549 slots (Q2_0 @512K) / 2574 slots (IQ3_S @512K); low-RAM mode **off**; CPU workers **engine default** (not set by us)
- MTP / speculative decoding **on** (`--spec 4 --spec-min-p 0.5`); draft acceptance observed ≈68 % in one Q2_0 sample
- sampling **engine defaults** (we only set `reasoning_effort`); vision off; calibration off
*Workload*
- 5-prompt mix (verbatim below) + an 8-task quality probe; **3 runs** for IQ3_S, **1 recorded run** for Q2_0
- per-prompt `max_tokens` in brackets; `reasoning_effort=none` everywhere
- cache state: expert-cache hit rate is logged per request — over **309 logged requests** (both quantisations) it ranged **26.5 %–87.0 %**, median **80.5 %**
(reproduce with: `grep -hoE "expert cache [0-9.]+% hit" <server log>`)
**Five-prompt mix — exact prompts (verbatim, `max_tokens` in brackets)**
1. `用一句话解释什么是混合专家模型(MoE)。` [128]
2. `写一段 200 字左右的短文,介绍咖啡的历史。` [400]
3. `用 Python 写快速排序,带注释,只给代码。` [500]
4. `A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost? Explain briefly.` [300]
5. `写一篇 600 字的技术短文,主题:为什么量化模型会损失精度。` [900]
| decode (comparison deliberately pairs a weighted mean against a median) | Q2_0 — 1 run, weighted mean | **IQ3_S — median of 3 runs** |
|---|---|---|
| tok/s | 45.0 (1015 tok / 22.6 s) | **33.7** (runs 33.1 / 33.7 / 34.9, range 33.1–34.9) |
| prefill, short prompts | 56–63 tok/s | 38–43 tok/s |
| expert cache | **5549 slots @512K** — measured with KV streamed to 2.58 GiB pinned RAM (the "5857" quoted elsewhere in our notes is a *different* configuration: 64K, KV fully in RAM) | **2574 slots @512K** |
| RAM after load | 54 GB | 68 GB |
**Long context (Q2_0 only)**
- prefill **1452 tok/s at 144,498 prompt tokens**; **855 tok/s at 510,900 prompt tokens**
- 512K needle test: **pass** — 510,900-token prompt, prefill **578 s**, correct answer
(engine log: `32768 of 524288 cells per QSA layer in VRAM, the K/V in 2.58 GiB of pinned RAM`)
- 512K vs 64K A/B on the 8-task probe: **the 7 tasks with a unique answer were byte-identical** ⇒ no visible yarn penalty for short prompts
(the 8th task is free-form writing and therefore not comparable)
**Quality probe — 8 short tasks: arithmetic `1234×5678`, bat-and-ball, `stressed` reversed, 17-sheep trick, weekday reasoning, strict JSON, short needle (60.3), constrained writing (≤100 chars, no "雨", must contain "铁")**
- Q2_0: 7/8 — missed bat-and-ball (answered 0.10); the probe scored 7/8 in **both** the 64K and the 512K configuration
- **IQ3_S: 8/8** — answered 0.05
**Observation**: IQ3_S is **~25% slower** here (33.7 vs 45.0) and its expert cache halves (5549 → 2574 slots) because
3.5-bit experts are bigger, so more experts are read from RAM; the small quality probe went the other way.
**One more thing (offered as a hook):** we have **reproduced issue #75 locally**. Under a long-context, tool-heavy task
(~136K prompt tokens) the model re-issued **one identical `bash` call 23 times in a row**; every `tool/call` had a matching
`tool/result` with `isError:false` and correct content, so the loop was **model-side**, not tooling. Happy to open a
separate report with the full logs if that is useful.
**A pattern we would like to hand over (explicitly labelled a hypothesis — 3 review sessions, one machine, no visibility inside the model)**
While producing the numbers above we ran three reviews of the *same* document with the *same* model and quantisation:
| session | prompt | outcome |
|---|---|---|
| 1 | “verify it yourself: grep the files, read `setup.py`, …” | **looped** — turn 3: 6 identical `bash` calls → our guard cancelled the turn; the model then looped again in **turn 4 (23 more identical calls)** — a de-dup bug in our guard let that turn through, and we killed it by hand |
| 2 | open-ended (“take a look at this draft”) | **clean** — finished quickly, still verifying ~12 facts on its own |
| 3 | **the same open-ended prompt as session 2** | **looped again** — one identical `bash` call 6 times, then our guard cancelled the turn automatically |
So the prompt style is **not** the deciding factor: the *same neutral prompt* produced both a clean run and a loop.
What is consistent is **where** it loops — in both loops the model was re-checking **one single fact in our draft**
(loop 1: “which header defines `__GLIBC_USE_IEC_60559_FUNCS_EXT_C23`, and the `find_nvcc` line numbers”;
loop 3: “confirm the long-context section contains no IQ3_S entry”), i.e. *self-verification that never reaches a stop
condition*. We would describe this as a possible **“self-verification → rationalisation → degradation”** tendency —
offered as a user-side observation, **not** as a finding, and n=3 sessions on one machine.
**One detail that may be directly useful for #75:** in session 3 the *harness itself* injected a warning after the 5th
identical call — “Repeated tool call detected … Do not call this tool with these exact arguments again. Inspect the latest
result and choose a different action, different arguments, or finish the task if enough evidence has been gathered.”
The model **ignored it and issued the same call again**. What actually ended the loop was a **hard cancel** of the turn
(our guard cancels on the 6th consecutive identical call). If that matches other reports, a stop *message* alone is not
sufficient — the loop ends when the turn is actually aborted. Happy to run controlled A/B prompts and publish raw logs.No site
Links install, modelos, releases.