Issues / #1300
#1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
open · @ZackO2o · 0 Kommentare · Auf GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Beschreibung
**What this is:** a two-card **V100-PCIE-32GB** Linux source build (sm_70, CUDA 12) running Qwen3.8-Flash-Next **IQ3_S** at the full **524,288-token** context with vision on. The 0.1.40 notes asked for Volta results; the existing Volta community reports are 16 GB cards or mixed Volta+Pascal pairs, so a 32 GB pair at 512K may be a useful data point.
The headline is not a speed number — it is that **decode tok/s on this engine is governed by tokens per verify window, which is a property of the content, not of the machine**, and that the engine can now show you this directly with the profiler added in 0.1.40.
## Hardware
- 2× **Tesla V100-PCIE-32GB** (sm_70), passive, server-cooled. GPU0 ~29.5 GiB used, GPU1 ~32.2 GiB used (of 32,768 MiB each).
- `nvidia-smi topo -m`: **GPU0↔GPU1 = SYS** — different CPU sockets, no P2P, no NVLink. Layer split (`--layer-split`) is the working path, as the docs describe.
- CPU: 2× Xeon E5-2673 v3 (12 cores each, 48 threads total, **AVX2 only** — no AVX-512). PCIe **Gen3 x16**.
- Memory: 125 GiB, **DDR3-1600** (see the note at the end — this is a real shortcoming of this box and it turned out not to matter).
- Linux, self-built engine with `-DCMAKE_CUDA_ARCHITECTURES=70 -DSTRATA_EXPERIMENTAL_SM60=ON`, CUDA 12.x.
## Configuration
`--max-context 524288 --rope-scaling yarn --rope-scale 2 --kv int8 --kv-resident 20480 --spec 4 --spec-min-p 0.70 --ple-io mmap --layer-split 20 --vision --vram-reserve-mib 700`
Two notes on the reasoning, both measured rather than assumed:
- **`--ple-io mmap`, not `direct` or `ram`.** `ram` is structurally wrong here, not just slower: locking the 26.8 GiB PLE table drives `MemFree` to 603 MB and the kernel then reclaims the process's *own* anonymous pages (`VmSwap` 1.45 GiB, `pswpin` +59%, `compact_stall` 324,567). The symptom is nasty because it does not look like a performance problem — the engine's own decode rate stays normal while a client sees tens of seconds of wall clock. `direct` pays an SSD read on every token's table lookup. `mmap` maps it into the page cache without locking, so rows in use stay resident and the rest stays reclaimable (`VmLck 0`, `VmSwap 0`). **We are not quoting a percentage for mmap over direct** — the arms were measured in separate windows and cannot be separated from run-to-run spread.
- **512K is configured and works**, with `--kv-resident` as a hard requirement rather than an optimisation. Without it the KV state takes the expert slots.
## Decode: the profiler explains what moves the number
Setting `STRATA_DECODE_TIMING=1` (with `STRATA_VERIFY_PROFILE=1`) logs a per-window breakdown per request. This is a genuinely useful addition — it is what let us stop guessing:
```
strata decode timing: 281 windows, avg T 2.04, 1.82 tokens/window, 25.80 ms/window
= verify 23.56 (GPU-reach wait 0.00 + per-layer host 0.00 [plan 0.05 actq 0.07 jobs 0.00 CPU 0.51]
+ stage 0.06) + commit/emit 0.37 + draft 1.59
per layer-window: CPU experts 0.10 (0.13 entries), VRAM hits 11.74, PCIe 0.01
```
Across seven profiled runs:
| tokens/window | ms/window | **ms/token** | CPU experts | VRAM hits |
| --: | --: | --: | --: | --: |
| 3.09 | 37.89 | **12.26** | 0.50 | 18.99 |
| 2.79 | 33.49 | **12.00** | 0.41 | 17.08 |
| 2.55 | 33.19 | **13.02** | 0.35 | 16.46 |
| 1.84 | 25.80 | **14.02** | 0.20 | 12.18 |
| 1.59 | 24.06 | **15.13** | 0.22 | 9.86 |
| 1.50 | 23.58 | **15.72** | 0.12 | 9.41 |
**ms/token is flat (12.0–15.7 ms) and ms/window is sub-linear in tokens/window** — about 23 ms fixed per window plus ~9 ms for each token the verifier accepts. So throughput is set by *tokens per window*, which is how many draft tokens the verifier accepts on this content. The CPU expert pool is 0.5–4% of window time and PCIe is 0.01 ms: on this box neither is close to being the bottleneck, even at 512K with a DDR3-1600 CPU half.
## Decode by content type — interleaved
Because a single number is not reproducible across content, here is the number with its content attached. **12 rounds per type, three types interleaved within one service lifetime (A,B,C,A,B,C…), first 2 rounds discarded, medians reported, and the whole thing run twice on separate occasions:**
| Content | Round 1 (median) | Round 2 (median) | Acceptance |
| --: | --: | --: | --: |
| prose (zh) | **84.8** | **80.8** | 78.8% |
| code | **82.8** | **79.3** | 83.0% |
| prose (en) | **70.9** | **68.3** | 81.5% |
- Content-to-content spread is **1.2×**; the same content across the two rounds differs by ≤5%, and the ordering reproduced.
- **Acceptance rate does not predict tok/s** — the lowest-acceptance arm (Chinese prose, 78.8%) was the *fastest*. What matters is tokens/window, i.e. acceptance combined with draft depth, not acceptance alone.
- The ordering is not the intuitive one (we expected code to win). Worth measuring per workload rather than assuming.
## Prefill (cold, engine counters)
| Prompt | Tokens read | Engine prompt time | Throughput |
| --: | --: | --: | --: |
| ~256 | 315 | 1,736 ms | 181.5 tok/s |
| ~1K | 1,004 | 1,892 ms | 530.8 tok/s |
| ~4K | 3,707 | 3,425 ms | 1,082.3 tok/s |
| ~16K | 14,572 | 8,653 ms | **1,684.1 tok/s** |
Fresh prompt per row — a repeat is served from the conversation cache and reports a meaningless figure. Above 16K we could not measure cold without changing the configuration, so we do not claim those rows.
## Other things that may be worth passing upstream
- **`--spec 4` beat `--spec 8`** on this box (79.4 vs 72.9 tok/s at 256K, acceptance 76–78% vs 61–65%) — shallow drafts, higher hit rate. The `--spec-min-p 0.70` row matched your calibration guidance.
- **`--batch` was a net loss here** and the reason is in `docs/BATCHING.md`: batch windows carry no MTP drafts, and this configuration depends on MTP heavily. Four concurrent requests queued at 0.90× — consistent with "one request at a time".
- **`--lookup-chain` (opt-in, default 0)**: we measured this as a **net loss on prose and code prompts** (73.7 vs 77.0). Reading the help, it drafts by matching repeated context, so it should only pay on templated/repetitive content — which neither of our benchmark prompts is. Not a bug report, just a note that "no measurable gain" here is expected rather than surprising, and it might be worth a line in the help saying which workloads it targets.
- **`--pool-affinity` has no "restrict to one NUMA node" mode** (only `all`/`auto`/`p-cores`). Wrapping the service in `numactl --cpunodebind=0` was **worse** (49–69 tok/s), so we left the defaults alone. Noting it because "stop the workers crossing QPI" looks like an obvious win on a two-socket box and there is no supported way to try it.
## What we did not measure
- **Full 512K recall.** Our deepest retrieval test is 241,255 tokens (correct). 512K is configured and answers, but recall at the very top of the window is untested.
- **Quality benchmarks.** No GSM8K/HumanEval/MMLU. Decode and prefill numbers here are safe to quote; quality numbers are absent.
- **A decode cost for 512K.** We measured 256K and 512K and the difference sits inside the run-to-run spread, so we make no claim either way.
- **Long-run stability.** These are figures from a live server, not a multi-day soak.
## A note on the DDR3-1600 finding
This box runs **DDR3-1600** although the E5-2673 v3 supports DDR4-2133, and `docs/` warns that RAM below its rated speed slows the CPU half. It is a real shortcoming, and the profiler is what showed it does not matter here: the CPU expert pool is 0.5–4% of window time, so even doubling memory bandwidth has a few percent of headroom. Recording it because "check your RAM speed" is the intuitive first move on a box like this and the profiler says it is not where the time goes.
Full configuration, build notes and the complete negative-results log (including several of our own earlier claims that we had to withdraw) are in our write-up: https://github.com/ZackO2o/Strata-Qwen3.8-Flash-Next-512k-v100x2
Happy to re-run anything on request — the box is up and the profiler is one environment variable away.
Mehr auf der Site
Links zu Install, Modellen, Releases.