Issues / #1079

#1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40

closed · @aleesposito85 · 2 Kommentare · Auf GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

**What this is:** a two-card Volta + Pascal Linux source build, running Qwen3.8-Flash-Next **IQ3_XXS** at the full **262,144-token** context with vision on and the P100 16 GB as a helper expert cache. The 0.1.40 notes asked for results from Pascal and Volta CUDA 12 engines; this covers Volta (sm_70) and a mixed Volta + Pascal pair, before and after 0.1.40.

## Headline

Same box, same two-card helper config, same bench — two runs each plus a confirm run:

| | 0.1.39 | 0.1.40 |
|---|---:|---:|
| Decode, 512-token answer | 65.3 / 65.2 (confirm 66.4) | **70.6 / 72.2** (confirm 70.7) tok/s |
| Decode, 128-token answers | 78.6-80.2 (two runs) | 79.6-80.3, then 84.7-84.9 (two runs) tok/s |
| Prompt, 13,561 fresh tokens | 444.0 / 444.5 | 443.5 / 444.0 tok/s |
| Quality battery | 25/25 | 25/25 |

That is **+6-11% decode** on the 512-token runs, prompt unchanged, answers equal. MTP acceptance was the same in both (~2/3 of drafts), so the gain is the expert path — #854 removes the helper's experts from the PCIe share (upstream measured 40 -> 66 on their rig; we see +6-11% on ours).

## Hardware

- Lenovo System x3650 M5 (2U server), Proxmox VE host; the engine runs in a Debian 12 LXC with the GPUs passed through.
- **GPU 0: Tesla V100-PCIE-32GB (sm_70)**, **GPU 1: Tesla P100-PCIE-16GB (sm_60)**. Both passive, server-cooled; both at their 250 W limits.
- **PCIe: x8 links on both cards** (idle `lspci` readout: 2.5 GT/s, width x8 "downgraded" from x16; not re-read under load), on **different CPU sockets** — `nvidia-smi topo -m` reports **SYS**, so no P2P. Strata's PCIe probe measured **1.6 GB/s** host->device -> `pcie_frac 0.04` (the 0.55 default would be wrong here; the probe sizing handled it).
- CPU: 2x Xeon E5-2699 v3 (18 cores each), 36 vCPUs to the container, **AVX2 only**. RAM: 120 GB, about 47 GB in use while serving (includes the file cache). Model files on a **SAS RAID5** (rotational) — hence `--ple-io ram` is kept on.
- Measured power draw with both GPUs inferring: **peak 553 W DC** (the box has 2x 750 W PSUs). Temperatures: in the layer-split runs the P100 hit 56 C while the V100 held near 70 C (the fan controller, which reads the max of both cards, ramped to keep it there); in the current helper setup the V100 peaks around 61 C and the P100 around 41 C. Day maximum fan reading: 77 C.

## Software

- Debian 12 container, Proxmox kernel 6.14.11-9-pve; NVIDIA driver **580.95.05**; CUDA toolkit **12.8.61** (nvcc), host compiler **GCC 12.2**.
- Strata **v0.1.40** (commit `1735d64`), **source build** of the CUDA 12 engine: `-DSTRATA_ENABLE_CUDA=ON -DSTRATA_EXPERIMENTAL_SM60=ON -DCMAKE_CUDA_ARCHITECTURES="70;60"` (58x sm_60 + 58x sm_70 cubins). 0.1.39 was built with the same recipe.
- **Build note for other GCC 12 users:** the v0.1.40 tree as tagged does not compile here — `include/strata/core/vmm.hpp:36` uses an unqualified `size_t` and the header only includes `<cstdint>`/`<vector>`, which GCC 12.2 rejects. One-token fix (`(std::size_t)`); filed as **#1076**. (Separate and also open: **#1071**, toolkits older than 12.5 in the same new file.)

## Model

- `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, **IQ3_XXS**, 2 shards (43.8 + 26.8 GiB, about 70.6 GiB total), served from a **Strata pack** (dense 1.5 GB, index, native expert list, tokenizer).
- MTP draft: the repo's **Q2_0 draft layer** (`mtp/rt`, 835 MiB VRAM per the engine log).
- Vision: **on** — `mmproj-Qwen3.8-Flash-Next-BF16` (907 MB) with the encoder from `tools/vision`, pinned to `cuda_device 0` (the V100) so it cannot land on the P100.

## Settings

```
strata ... --pack <pack> --native shard1 --ple-gguf shard2
  --max-context 262144 --kv int8 --kv-resident 32768
  --expert-cache auto                # 15,354 slots / 24.90 GiB on the V100
  --expert-cache-device1 auto        # 9,189 experts / 15.01 GiB on the P100
  --remote-expert-opt
  --spec 4 --spec-min-p 0.5 --mtp <mtp/rt>
  --ple-io ram --vram-reserve-mib 700
  --vision
env: CUDA_VISIBLE_DEVICES=0,1
```

Decode cache hit rate reported: **100.0% warm on both versions** (early requests while the cache fills are lower; e.g. 87.8-98.8% on 0.1.39 during warm-up). Restart to ready: **135-136 s** on 0.1.40 with a warm file cache (both observed runs; earlier 0.1.39 restarts on the same box measured 105 s and 180 s, so treat that as cache-state noise rather than a claim about the read-ahead). KV for 262K: 32,768 of each QSA layer's cells in VRAM, the rest (3.09 GiB) in pinned RAM.

## How it was measured

A custom bench script (two runs per version plus a confirm run): a warm smoke request, four short probes, three 128-token writes, one 512-token write, one 13,561-token prompt — all fresh (`0 reused + 13561 read`, per the engine log), greedy, server-reported numbers. The quality battery is 25 deterministic auto-graded probes (arithmetic, word problems, letter counting, logic, code): **25/25 on both versions, no probe changed**.

Not the repo's `benchmark.py`; 2-3 runs per config rather than 3+. One client, one request at a time (the server serves one sequence).

## Two more data points for the multi-GPU work

- **Layer split vs helper on this pair (0.1.39):** `"gpu":[0,1]` + `layer_split auto` placed 47 of 48 layers on the V100 and starved the prompt path — prefill fell to **188-194 tok/s (-57%)**; a balanced K=28 split restored prefill (431-440) but cost decode (**41-44, -26%**). The helper cache won both ways. The engine log itself suggested the helper mode "for a card this small".
- **IQ3_S vs IQ3_XXS with the helper cache:** same quality (25/25 both), but IQ3_S costs **prefill -25%** (334 vs 444) and **decode -7%** (59-61 vs 65) — its bigger expert arena fits fewer slots, so more experts cross the link. XXS stayed.

Limitations: one machine, one model size; my own bench script rather than the repo's; the two versions were measured same-box/same-day but not interleaved; both sides ran the identical two-card helper configuration.

Happy to run anything specific if a Volta test would help (the 0.1.40 "testers wanted" list, or the `--pipeline-windows` / resident-split soak — the layer-split path is the one we don't use day to day).

Mehr auf der Site

Links zu Install, Modellen, Releases.