反馈 / #1639

#1639 Community benchmark: 5x Tesla P100 (sm_60), Flash-Next IQ3_S, 262K context — STRATA_STAGE_TRIM data + a Pascal Q6_K kernel

open · @thedrzinger · 0 评论 · 去 GitHub 看

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindows

说明

Measured on 2026-10-08/09 on five Tesla P100s with Flash-Next IQ3_S at the native 262,144-token context. Main limitations: **2 runs per configuration** in the main comparison (1 run in the context sweep), and the machine is a Proxmox LXC container with 24 CPU threads and 52 GB of RAM. **Only the baseline (A0) used a stock engine.** Every other number comes from a locally modified build of `fb58e0d`, described under "Hardware and software".

Two findings that may be useful beyond Pascal:

1. **`STRATA_STAGE_TRIM=1` with explicit split points made the largest difference on this rig.** With 5 cards, every card kept a full dense copy. With the trim, the expert caches went from 19,364 to 23,702 of 24,576 profiled experts (~98.0% → ~99.9% of the routed mass), and CUDA1/CUDA3 held all of theirs. Measured together with a context change (see table A).
2. **On Pascal, the Q6_K decode MMVQ is far below memory bandwidth.**
   - **Cause:** 2-byte loads on the 210-byte blocks, plus `dp4a` emulated per column. I measured 2-byte loads topping out at ~360 GB/s and 8-byte loads at ~600 GB/s on a P100.
   - **The kernel:** I wrote a Pascal-only kernel that reads the existing `native_q6_k_pack` planes with 8-byte loads and does FP32 math on exact small integers. It's 1.1-2.2x faster per call and was +7-20% on decode here (table A). It's not bitwise equal to the dp4a kernels (relative difference ~2e-7), so as a PR it would be opt-in.
   - **Offer:** happy to open it as a separate opt-in PR if that's wanted.

## Hardware and software

- **GPUs:** 1x Tesla P100-PCIE-16GB + 4x Tesla P100-PCIE-12GB (64 GB VRAM total). No NVLink, two PCIe root complexes (GPU0-1 on NUMA node 0, GPU2-4 on node 1). All PCIe 3.0 x16 capable; the engine's probe measured 11.8 / 10.4 / 6.2 / 10.5 / 10.4 GB/s host→device. 200 W power limit each.
- **CPU, RAM, storage:** 2x Xeon E5-2650 v4 (AVX2, no AVX-512); the container gets 24 threads, and the engine used 17 expert-pool workers + the host thread. 52 GB RAM in the container, NVMe storage.
- **OS:** Ubuntu 26.04 LXC on Proxmox (kernel 7.0.14-12-pve), NVIDIA driver 580.167.08, CUDA 12.9 (nvcc 12.9, host compiler g++-14).
- **Strata:** `fb58e0d` (engine 0.1.41), built from source with `STRATA_EXPERIMENTAL_SM60=1` and `-DCMAKE_CUDA_ARCHITECTURES=60`. Two builds were used:
  - **Stock:** the engine `setup.sh --build --cuda 12` compiled from unmodified `fb58e0d`. Used for config A0 only.
  - **Modified:** `fb58e0d` plus my local changes, not upstream. Used for A1, A2 and everything after (sections B and C, needles, HumanEval+, vision). The changes:
    - the Pascal Q6_K kernel, switchable with an env var for the A1/A2 comparison;
    - the bit-exact PRMT/XMAD `dp4a` emulation in `include/strata/kernels/dp4a.hpp`;
    - the same sequence patched into ggml's sm_60 `dp4a` fallback (prompt-path MMQ).

    With the kernel off (A1), the only difference from stock is the `dp4a` change, which is bit-exact. The Python server, setup and everything else are unmodified.
- **Background workloads:** none on the GPUs during the runs.

## Model and configuration

- **Model:** `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, IQ3_S (both shards), prepared by the unmodified installer.
- **Bundled files:** expert profile, MTP draft head from the installer (`mtp/rt`).
- **Memory mode:** low-RAM resident mode (`--resident-experts`), chosen by setup for 52 GB of RAM.
- **Engine settings:** int8 KV, `--expert-cache auto`, `--prefill auto`, `--spec 4 --spec-min-p 0.5`. No calibration, experimental speed projection off.
- **Setup command:**

```text
STRATA_EXPERIMENTAL_SM60=1 ./setup.sh --yes --family qwen --model IQ3_S --gguf-dir <IQ3_S dir> \
  --gpus 0,1,2,3,4 --cuda 12 --build --vision no --no-start
```

Configurations compared (same machine, same workload, back to back):

| Config | What changed |
|---|---|
| **A0** setup defaults | **stock engine**, config as written by setup: 32K context, `layer_split auto` (picked K=10,19,29,38) |
| **A1** + trim | **modified engine** with the Pascal Q6_K kernel **off**, `STRATA_STAGE_TRIM=1`, `"layer_split": "10,19,29,38"`, `--max-context 262144` |
| **A2** + Pascal kernel | A1 with the Pascal Q6_K kernel on (**modified engine**) |

A0 → A1 changes only stock settings (trim, split points, context). The engine difference is the bit-exact `dp4a` change, which measured no decode difference in isolation. A1 → A2 is the Pascal kernel alone.

## Method

- **Workload:** `bench/results/2026-10-06-community-2x-p100/benchmark.py`, unchanged. Its synthetic code-explanation prompts are fresh for every request (0 reused tokens, checked in the engine log).
- **Generation:** greedy, reasoning off, 256-token cap. Every run hit the cap.
- **Timing:** engine-reported prompt and decode rates; TTFT measured at the client over loopback.
- **Cache state:** one warm-up request was excluded, and the expert cache stayed warm between runs.
- **Restarts:** the engine was restarted between configurations, not between runs.

## Results

### A. Settings comparison (2 runs each; median [range])

| Config | Prompt tokens | Prompt tok/s | Decode tok/s | TTFT s |
| --- | ---: | --- | --- | --- |
| A0 setup defaults | 511 | 93.2 [83.5–102.8] | 31.4 [29.7–33.2] | 5.59 [5.01–6.17] |
| A1 + trim, 262K | 511 | 103.8 [89.4–118.2] | 34.1 [33.2–34.9] | 5.07 [4.37–5.77] |
| A2 + Pascal kernel | 511 | 104.5 [90.9–118.2] | **37.2** [36.7–37.7] | 5.01 [4.37–5.66] |
| A0 setup defaults | 3,999 | 234.4 [234.0–234.8] | 35.8 [35.0–36.5] | 17.11 [17.08–17.14] |
| A1 + trim, 262K | 3,999 | 255.2 [254.6–255.9] | 36.5 [34.7–38.4] | 15.73 [15.68–15.78] |
| A2 + Pascal kernel | 3,999 | 256.7 [255.5–257.8] | **38.9** [37.7–40.1] | 15.63 [15.56–15.71] |
| A0 setup defaults | 16,000 | 401.3 [368.0–434.7] | 32.7 [31.5–33.8] | 40.22 [36.87–43.57] |
| A1 + trim, 262K | 16,000 | 467.4 [463.7–471.2] | 33.2 [33.0–33.4] | 34.32 [34.05–34.59] |
| A2 + Pascal kernel | 16,000 | 467.0 [461.6–472.4] | **40.1** [37.6–42.6] | 34.34 [33.94–34.74] |

Decode hit rate 97.6–99.8% in all runs; drafts accepted 142–165 of 203–241 per run.

Expert slots by card (CUDA0..4) and the profiled routed mass the caches held:

| | Slots, CUDA0 / 1 / 2 / 3 / 4 | Routed mass held |
|---|---|---|
| A0 (no trim) | 5120 / 3916 / 3678 / 3738 / 2749 | ~98.0% |
| A1 / A2 (trim) | 5120 / 4608 / 4956 / 4608 / 3736 | ~99.9% |

At 262K context the slots are a little lower: 4531 on CUDA2, 3232 on CUDA4.

### B. Context sweep (config A2; 1 run each)

| Context | Decode tok/s, 3,999 / 64,000-token prompt | Prompt tok/s, 64,000 | Notes |
|---|---|---|---|
| 131,072 | 38.6 / 39.3 | 841.7 | |
| 196,608 | 38.2 / 39.7 | 802.0 | |
| 262,144 | 37.9 / 37.4 | 796.6 | 200,000-token prompt: decode **35.9**, prompt 958.8 tok/s, TTFT 209 s, hit 99.5% |
| 262,144, `--kv-resident 32768` | 38.3 / 39.1 | 789.1 | ~400 more experts cached; no clear difference |
| 524,288, YaRN x2 | 38.9 / 41.2 | 795.3 | experimental, quality not checked |

### C. The Pascal kernel in isolation

Per call, measured against `native_mmvq(14, ...)` on identical data. Relative max difference 1-2e-7 for 1-8 columns.

| Shape | Columns | Stock | Pascal kernel | Speedup |
|---|---:|---:|---:|---:|
| 2560x12288 | 1 / 4 / 8 | 135 / 255 / 434 µs | 80 / 139 / 208 µs | 1.70x / 1.84x / 2.08x |
| 4096x2560 | 1 / 4 / 8 | 47 / 88 / 144 µs | 35 / 60 / 87 µs | 1.35x / 1.47x / 1.67x |
| head 2560x248320 | 1 / 4 / 8 | 2.47 / 4.73 / 7.91 ms | 1.26 / 2.30 / 3.64 ms | 1.97x / 2.06x / 2.17x |

Memory at A2, 262K, with GPU vision on GPU0:
- **VRAM used:** 13.1 / 10.7 / 11.6 / 11.2 / 12.0 GiB.
- **RAM:** ~11-15 GB used (5.1 GiB of resident experts).

No OOM, no paging.

## Correctness and limitations

- **Needles** (`tools/needle_bench.py`, 262K context): 6 of 6 found at 8K (8,064 tokens) and 190K (189,451 tokens), depths 10/50/90.
- **HumanEval+** (EvalPlus 0.1.10, greedy, thinking off, standard EvalPlus prompt, 164 problems):

  | Run | HumanEval | HumanEval+ |
  |---|---|---|
  | Kernel on, run 1 | 96.3% | 94.5% |
  | Kernel on, run 2 | 96.3% | 92.1% |
  | Kernel off | 97.6% | 95.1% |

  Two kernel-on runs already differ on 6 problems (cache and CPU/GPU rounding between restarts), and kernel on vs off differ on 2-5. So no measurable effect from the kernel.
- **Vision** (GPU, encoder pinned to GPU0 via `"cuda_device": 0`): read the text and shapes of a test image correctly. Expert caches and decode speed were unchanged.
- **Not tested:**
  - other quants (Q2_0, IQ2_XS, IQ3_XXS) and other models (Swift 1.5, Coder);
  - fewer than 5 GPUs;
  - `--pipeline-windows`;
  - prompts past 200K tokens;
  - 3+ runs per configuration;
  - a cold page cache.

  A0 differs from A1 in context (32K vs 262K) as well as the trim, so table A does not isolate the trim alone. The slot counts do.
- **Other gains:** the stock dp4a emulation could also use PRMT + 16-bit MADs (bit-exact, 1.46x in isolation). It had no measurable effect on decode here, so I'm not proposing it.

本站相关内容

相关页面的快捷入口。