Pull requests / #1452
#1452 bench: community report, Tesla V100 32 GB + P100 16 GB expert-helper (mixed Volta/Pascal), Flash-Next IQ3_XXS
open · @aleesposito85 · 0 comentários · No GitHub
BenchmarksNVIDIA / CUDAModels & quantsDocumentation
Descrição
Resolves #1079 (data report; the issue itself was closed as answered - this keeps the numbers with the other community reports). ## Summary Community benchmark report: **Tesla V100-PCIE-32GB (sm_70) + Tesla P100-PCIE-16GB (sm_60) as the expert-helper cache** (`--expert-cache-device1 auto --remote-expert-opt`), Qwen3.8-Flash-Next IQ3_XXS, 262,144-token context, engine 0.1.40.2 built from source (CUDA 12.8, sm_70+sm_60). This is the first mixed Volta/Pascal report: the maintainer's reply on #1079 asked for a bench PR with the data. ## What changed Adds `bench/results/2026-10-07-community-v100-p100-helper/`: - `README.md` - report in the format of docs/COMMUNITY_BENCHMARKS.md (hardware, software, model hashes, exact launch config, method, results table, correctness checks, limitations). - `results.json` - per-run engine timing blocks for all measured runs. - `engine-timings.txt` - the engine's own timing lines for the battery. - `needles.json` - `tools/needle_bench.py` output: 6 of 6 found (32k x3 depths, 128k x3 depths). - `strata-iq3_xxs.json` - the complete server config. - `model-sha256.txt` - model, mmproj, MTP and expert-profile hashes. - `bench_report.py` - the measurement script (runs unchanged against any Strata server). Key numbers (warm, single-sequence, 256-token cap, greedy, reasoning off): prompt 1,436.9 tok/s median at 17,971 tokens; 1,446.4 at 44,011; 1,410.5 single run at 119,731. Decode ~73 tok/s median. Two configuration notes measured for this rig: `--prefill auto:32768` (the default 8192 chunk cap costs 50-100% prompt speed on helper-cache rigs - each >=1024-token chunk streams every non-resident expert) and `STRATA_PREFILL_CPU_SHARE=auto` (-23% TTFT on sub-1024-token prompts, output-identical in a fixed-prompt check). No engine changes in this PR - results only. ## Extra Notes The rig has no P2P between the cards (different CPU sockets); decode includes the helper round-trip. The 0.1.39-to-0.1.40 decode delta and the helper-cache A/B were reported earlier in #1079.
No site
Links install, modelos, releases.