Issues / #1639
#1639 Community benchmark: 5x Tesla P100 (sm_60), Flash-Next IQ3_S, 262K context — STRATA_STAGE_TRIM data + a Pascal Q6_K kernel
open · @thedrzinger · 0 comments · View on GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindows
Description
Measured on 2026-10-08/09 on five Tesla P100s with Flash-Next IQ3_S at the native 262,144-token context. Main limitations: **2 runs per configuration** in the main comparison (1 run in the context sweep), and the machine is a Proxmox LXC container with 24 CPU threads and 52 GB of RAM. **Only the baseline (A0) used a stock engine.** Every other number comes from a locally modified build of `fb58e0d`, described under "Hardware and software".
Two findings that may be useful beyond Pascal:
1. **`STRATA_STAGE_TRIM=1` with explicit split points made the largest difference on this rig.** With 5 cards, every card kept a full dense copy. With the trim, the expert caches went from 19,364 to 23,702 of 24,576 profiled experts (~98.0% → ~99.9% of the routed mass), and CUDA1/CUDA3 held all of theirs. Measured together with a context change (see table A).
2. **On Pascal, the Q6_K decode MMVQ is far below memory bandwidth.**
- **Cause:** 2-byte loads on the 210-byte blocks, plus `dp4a` emulated per column. I measured 2-byte loads topping out at ~360 GB/s and 8-byte loads at ~600 GB/s on a P100.
- **The kernel:** I wrote a Pascal-only kernel that reads the existing `native_q6_k_pack` planes with 8-byte loads and does FP32 math on exact small integers. It's 1.1-2.2x faster per call and was +7-20% on decode here (table A). It's not bitwise equal to the dp4a kernels (relative difference ~2e-7), so as a PR it would be opt-in.
- **Offer:** happy to open it as a separate opt-in PR if that's wanted.
## Hardware and software
- **GPUs:** 1x Tesla P100-PCIE-16GB + 4x Tesla P100-PCIE-12GB (64 GB VRAM total). No NVLink, two PCIe root complexes (GPU0-1 on NUMA node 0, GPU2-4 on node 1). All PCIe 3.0 x16 capable; the engine's probe measured 11.8 / 10.4 / 6.2 / 10.5 / 10.4 GB/s host→device. 200 W power limit each.
- **CPU, RAM, storage:** 2x Xeon E5-2650 v4 (AVX2, no AVX-512); the container gets 24 threads, and the engine used 17 expert-pool workers + the host thread. 52 GB RAM in the container, NVMe storage.
- **OS:** Ubuntu 26.04 LXC on Proxmox (kernel 7.0.14-12-pve), NVIDIA driver 580.167.08, CUDA 12.9 (nvcc 12.9, host compiler g++-14).
- **Strata:** `fb58e0d` (engine 0.1.41), built from source with `STRATA_EXPERIMENTAL_SM60=1` and `-DCMAKE_CUDA_ARCHITECTURES=60`. Two builds were used:
- **Stock:** the engine `setup.sh --build --cuda 12` compiled from unmodified `fb58e0d`. Used for config A0 only.
- **Modified:** `fb58e0d` plus my local changes, not upstream. Used for A1, A2 and everything after (sections B and C, needles, HumanEval+, vision). The changes:
- the Pascal Q6_K kernel, switchable with an env var for the A1/A2 comparison;
- the bit-exact PRMT/XMAD `dp4a` emulation in `include/strata/kernels/dp4a.hpp`;
- the same sequence patched into ggml's sm_60 `dp4a` fallback (prompt-path MMQ).
With the kernel off (A1), the only difference from stock is the `dp4a` change, which is bit-exact. The Python server, setup and everything else are unmodified.
- **Background workloads:** none on the GPUs during the runs.
## Model and configuration
- **Model:** `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, IQ3_S (both shards), prepared by the unmodified installer.
- **Bundled files:** expert profile, MTP draft head from the installer (`mtp/rt`).
- **Memory mode:** low-RAM resident mode (`--resident-experts`), chosen by setup for 52 GB of RAM.
- **Engine settings:** int8 KV, `--expert-cache auto`, `--prefill auto`, `--spec 4 --spec-min-p 0.5`. No calibration, experimental speed projection off.
- **Setup command:**
```text
STRATA_EXPERIMENTAL_SM60=1 ./setup.sh --yes --family qwen --model IQ3_S --gguf-dir <IQ3_S dir> \
--gpus 0,1,2,3,4 --cuda 12 --build --vision no --no-start
```
Configurations compared (same machine, same workload, back to back):
| Config | What changed |
|---|---|
| **A0** setup defaults | **stock engine**, config as written by setup: 32K context, `layer_split auto` (picked K=10,19,29,38) |
| **A1** + trim | **modified engine** with the Pascal Q6_K kernel **off**, `STRATA_STAGE_TRIM=1`, `"layer_split": "10,19,29,38"`, `--max-context 262144` |
| **A2** + Pascal kernel | A1 with the Pascal Q6_K kernel on (**modified engine**) |
A0 → A1 changes only stock settings (trim, split points, context). The engine difference is the bit-exact `dp4a` change, which measured no decode difference in isolation. A1 → A2 is the Pascal kernel alone.
## Method
- **Workload:** `bench/results/2026-10-06-community-2x-p100/benchmark.py`, unchanged. Its synthetic code-explanation prompts are fresh for every request (0 reused tokens, checked in the engine log).
- **Generation:** greedy, reasoning off, 256-token cap. Every run hit the cap.
- **Timing:** engine-reported prompt and decode rates; TTFT measured at the client over loopback.
- **Cache state:** one warm-up request was excluded, and the expert cache stayed warm between runs.
- **Restarts:** the engine was restarted between configurations, not between runs.
## Results
### A. Settings comparison (2 runs each; median [range])
| Config | Prompt tokens | Prompt tok/s | Decode tok/s | TTFT s |
| --- | ---: | --- | --- | --- |
| A0 setup defaults | 511 | 93.2 [83.5–102.8] | 31.4 [29.7–33.2] | 5.59 [5.01–6.17] |
| A1 + trim, 262K | 511 | 103.8 [89.4–118.2] | 34.1 [33.2–34.9] | 5.07 [4.37–5.77] |
| A2 + Pascal kernel | 511 | 104.5 [90.9–118.2] | **37.2** [36.7–37.7] | 5.01 [4.37–5.66] |
| A0 setup defaults | 3,999 | 234.4 [234.0–234.8] | 35.8 [35.0–36.5] | 17.11 [17.08–17.14] |
| A1 + trim, 262K | 3,999 | 255.2 [254.6–255.9] | 36.5 [34.7–38.4] | 15.73 [15.68–15.78] |
| A2 + Pascal kernel | 3,999 | 256.7 [255.5–257.8] | **38.9** [37.7–40.1] | 15.63 [15.56–15.71] |
| A0 setup defaults | 16,000 | 401.3 [368.0–434.7] | 32.7 [31.5–33.8] | 40.22 [36.87–43.57] |
| A1 + trim, 262K | 16,000 | 467.4 [463.7–471.2] | 33.2 [33.0–33.4] | 34.32 [34.05–34.59] |
| A2 + Pascal kernel | 16,000 | 467.0 [461.6–472.4] | **40.1** [37.6–42.6] | 34.34 [33.94–34.74] |
Decode hit rate 97.6–99.8% in all runs; drafts accepted 142–165 of 203–241 per run.
Expert slots by card (CUDA0..4) and the profiled routed mass the caches held:
| | Slots, CUDA0 / 1 / 2 / 3 / 4 | Routed mass held |
|---|---|---|
| A0 (no trim) | 5120 / 3916 / 3678 / 3738 / 2749 | ~98.0% |
| A1 / A2 (trim) | 5120 / 4608 / 4956 / 4608 / 3736 | ~99.9% |
At 262K context the slots are a little lower: 4531 on CUDA2, 3232 on CUDA4.
### B. Context sweep (config A2; 1 run each)
| Context | Decode tok/s, 3,999 / 64,000-token prompt | Prompt tok/s, 64,000 | Notes |
|---|---|---|---|
| 131,072 | 38.6 / 39.3 | 841.7 | |
| 196,608 | 38.2 / 39.7 | 802.0 | |
| 262,144 | 37.9 / 37.4 | 796.6 | 200,000-token prompt: decode **35.9**, prompt 958.8 tok/s, TTFT 209 s, hit 99.5% |
| 262,144, `--kv-resident 32768` | 38.3 / 39.1 | 789.1 | ~400 more experts cached; no clear difference |
| 524,288, YaRN x2 | 38.9 / 41.2 | 795.3 | experimental, quality not checked |
### C. The Pascal kernel in isolation
Per call, measured against `native_mmvq(14, ...)` on identical data. Relative max difference 1-2e-7 for 1-8 columns.
| Shape | Columns | Stock | Pascal kernel | Speedup |
|---|---:|---:|---:|---:|
| 2560x12288 | 1 / 4 / 8 | 135 / 255 / 434 µs | 80 / 139 / 208 µs | 1.70x / 1.84x / 2.08x |
| 4096x2560 | 1 / 4 / 8 | 47 / 88 / 144 µs | 35 / 60 / 87 µs | 1.35x / 1.47x / 1.67x |
| head 2560x248320 | 1 / 4 / 8 | 2.47 / 4.73 / 7.91 ms | 1.26 / 2.30 / 3.64 ms | 1.97x / 2.06x / 2.17x |
Memory at A2, 262K, with GPU vision on GPU0:
- **VRAM used:** 13.1 / 10.7 / 11.6 / 11.2 / 12.0 GiB.
- **RAM:** ~11-15 GB used (5.1 GiB of resident experts).
No OOM, no paging.
## Correctness and limitations
- **Needles** (`tools/needle_bench.py`, 262K context): 6 of 6 found at 8K (8,064 tokens) and 190K (189,451 tokens), depths 10/50/90.
- **HumanEval+** (EvalPlus 0.1.10, greedy, thinking off, standard EvalPlus prompt, 164 problems):
| Run | HumanEval | HumanEval+ |
|---|---|---|
| Kernel on, run 1 | 96.3% | 94.5% |
| Kernel on, run 2 | 96.3% | 92.1% |
| Kernel off | 97.6% | 95.1% |
Two kernel-on runs already differ on 6 problems (cache and CPU/GPU rounding between restarts), and kernel on vs off differ on 2-5. So no measurable effect from the kernel.
- **Vision** (GPU, encoder pinned to GPU0 via `"cuda_device": 0`): read the text and shapes of a test image correctly. Expert caches and decode speed were unchanged.
- **Not tested:**
- other quants (Q2_0, IQ2_XS, IQ3_XXS) and other models (Swift 1.5, Coder);
- fewer than 5 GPUs;
- `--pipeline-windows`;
- prompts past 200K tokens;
- 3+ runs per configuration;
- a cold page cache.
A0 differs from A1 in context (32K vs 262K) as well as the trim, so table A does not isolate the trim alone. The slot counts do.
- **Other gains:** the stock dp4a emulation could also use PRMT + 16-bit MADs (bit-exact, 1.46x in isolation). It had no measurable effect on decode here, so I'm not proposing it.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.