Issues / #1079
#1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40
closed · @aleesposito85 · 2 Kommentare · Auf GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
**What this is:** a two-card Volta + Pascal Linux source build, running Qwen3.8-Flash-Next **IQ3_XXS** at the full **262,144-token** context with vision on and the P100 16 GB as a helper expert cache. The 0.1.40 notes asked for results from Pascal and Volta CUDA 12 engines; this covers Volta (sm_70) and a mixed Volta + Pascal pair, before and after 0.1.40. ## Headline Same box, same two-card helper config, same bench — two runs each plus a confirm run: | | 0.1.39 | 0.1.40 | |---|---:|---:| | Decode, 512-token answer | 65.3 / 65.2 (confirm 66.4) | **70.6 / 72.2** (confirm 70.7) tok/s | | Decode, 128-token answers | 78.6-80.2 (two runs) | 79.6-80.3, then 84.7-84.9 (two runs) tok/s | | Prompt, 13,561 fresh tokens | 444.0 / 444.5 | 443.5 / 444.0 tok/s | | Quality battery | 25/25 | 25/25 | That is **+6-11% decode** on the 512-token runs, prompt unchanged, answers equal. MTP acceptance was the same in both (~2/3 of drafts), so the gain is the expert path — #854 removes the helper's experts from the PCIe share (upstream measured 40 -> 66 on their rig; we see +6-11% on ours). ## Hardware - Lenovo System x3650 M5 (2U server), Proxmox VE host; the engine runs in a Debian 12 LXC with the GPUs passed through. - **GPU 0: Tesla V100-PCIE-32GB (sm_70)**, **GPU 1: Tesla P100-PCIE-16GB (sm_60)**. Both passive, server-cooled; both at their 250 W limits. - **PCIe: x8 links on both cards** (idle `lspci` readout: 2.5 GT/s, width x8 "downgraded" from x16; not re-read under load), on **different CPU sockets** — `nvidia-smi topo -m` reports **SYS**, so no P2P. Strata's PCIe probe measured **1.6 GB/s** host->device -> `pcie_frac 0.04` (the 0.55 default would be wrong here; the probe sizing handled it). - CPU: 2x Xeon E5-2699 v3 (18 cores each), 36 vCPUs to the container, **AVX2 only**. RAM: 120 GB, about 47 GB in use while serving (includes the file cache). Model files on a **SAS RAID5** (rotational) — hence `--ple-io ram` is kept on. - Measured power draw with both GPUs inferring: **peak 553 W DC** (the box has 2x 750 W PSUs). Temperatures: in the layer-split runs the P100 hit 56 C while the V100 held near 70 C (the fan controller, which reads the max of both cards, ramped to keep it there); in the current helper setup the V100 peaks around 61 C and the P100 around 41 C. Day maximum fan reading: 77 C. ## Software - Debian 12 container, Proxmox kernel 6.14.11-9-pve; NVIDIA driver **580.95.05**; CUDA toolkit **12.8.61** (nvcc), host compiler **GCC 12.2**. - Strata **v0.1.40** (commit `1735d64`), **source build** of the CUDA 12 engine: `-DSTRATA_ENABLE_CUDA=ON -DSTRATA_EXPERIMENTAL_SM60=ON -DCMAKE_CUDA_ARCHITECTURES="70;60"` (58x sm_60 + 58x sm_70 cubins). 0.1.39 was built with the same recipe. - **Build note for other GCC 12 users:** the v0.1.40 tree as tagged does not compile here — `include/strata/core/vmm.hpp:36` uses an unqualified `size_t` and the header only includes `<cstdint>`/`<vector>`, which GCC 12.2 rejects. One-token fix (`(std::size_t)`); filed as **#1076**. (Separate and also open: **#1071**, toolkits older than 12.5 in the same new file.) ## Model - `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, **IQ3_XXS**, 2 shards (43.8 + 26.8 GiB, about 70.6 GiB total), served from a **Strata pack** (dense 1.5 GB, index, native expert list, tokenizer). - MTP draft: the repo's **Q2_0 draft layer** (`mtp/rt`, 835 MiB VRAM per the engine log). - Vision: **on** — `mmproj-Qwen3.8-Flash-Next-BF16` (907 MB) with the encoder from `tools/vision`, pinned to `cuda_device 0` (the V100) so it cannot land on the P100. ## Settings ``` strata ... --pack <pack> --native shard1 --ple-gguf shard2 --max-context 262144 --kv int8 --kv-resident 32768 --expert-cache auto # 15,354 slots / 24.90 GiB on the V100 --expert-cache-device1 auto # 9,189 experts / 15.01 GiB on the P100 --remote-expert-opt --spec 4 --spec-min-p 0.5 --mtp <mtp/rt> --ple-io ram --vram-reserve-mib 700 --vision env: CUDA_VISIBLE_DEVICES=0,1 ``` Decode cache hit rate reported: **100.0% warm on both versions** (early requests while the cache fills are lower; e.g. 87.8-98.8% on 0.1.39 during warm-up). Restart to ready: **135-136 s** on 0.1.40 with a warm file cache (both observed runs; earlier 0.1.39 restarts on the same box measured 105 s and 180 s, so treat that as cache-state noise rather than a claim about the read-ahead). KV for 262K: 32,768 of each QSA layer's cells in VRAM, the rest (3.09 GiB) in pinned RAM. ## How it was measured A custom bench script (two runs per version plus a confirm run): a warm smoke request, four short probes, three 128-token writes, one 512-token write, one 13,561-token prompt — all fresh (`0 reused + 13561 read`, per the engine log), greedy, server-reported numbers. The quality battery is 25 deterministic auto-graded probes (arithmetic, word problems, letter counting, logic, code): **25/25 on both versions, no probe changed**. Not the repo's `benchmark.py`; 2-3 runs per config rather than 3+. One client, one request at a time (the server serves one sequence). ## Two more data points for the multi-GPU work - **Layer split vs helper on this pair (0.1.39):** `"gpu":[0,1]` + `layer_split auto` placed 47 of 48 layers on the V100 and starved the prompt path — prefill fell to **188-194 tok/s (-57%)**; a balanced K=28 split restored prefill (431-440) but cost decode (**41-44, -26%**). The helper cache won both ways. The engine log itself suggested the helper mode "for a card this small". - **IQ3_S vs IQ3_XXS with the helper cache:** same quality (25/25 both), but IQ3_S costs **prefill -25%** (334 vs 444) and **decode -7%** (59-61 vs 65) — its bigger expert arena fits fewer slots, so more experts cross the link. XXS stayed. Limitations: one machine, one model size; my own bench script rather than the repo's; the two versions were measured same-box/same-day but not interleaved; both sides ran the identical two-card helper configuration. Happy to run anything specific if a Volta test would help (the 0.1.40 "testers wanted" list, or the `--pipeline-windows` / resident-split soak — the layer-split path is the one we don't use day to day).
Mehr auf der Site
Links zu Install, Modellen, Releases.