Pull requests / #1452

#1452 bench: community report, Tesla V100 32 GB + P100 16 GB expert-helper (mixed Volta/Pascal), Flash-Next IQ3_XXS

open · @aleesposito85 · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

描述

Resolves #1079 (data report; the issue itself was closed as answered - this keeps the numbers with the other community reports).

## Summary

Community benchmark report: **Tesla V100-PCIE-32GB (sm_70) + Tesla P100-PCIE-16GB (sm_60) as the expert-helper cache** (`--expert-cache-device1 auto --remote-expert-opt`), Qwen3.8-Flash-Next IQ3_XXS, 262,144-token context, engine 0.1.40.2 built from source (CUDA 12.8, sm_70+sm_60). This is the first mixed Volta/Pascal report: the maintainer's reply on #1079 asked for a bench PR with the data.

## What changed

Adds `bench/results/2026-10-07-community-v100-p100-helper/`:

- `README.md` - report in the format of docs/COMMUNITY_BENCHMARKS.md (hardware, software, model hashes, exact launch config, method, results table, correctness checks, limitations).
- `results.json` - per-run engine timing blocks for all measured runs.
- `engine-timings.txt` - the engine's own timing lines for the battery.
- `needles.json` - `tools/needle_bench.py` output: 6 of 6 found (32k x3 depths, 128k x3 depths).
- `strata-iq3_xxs.json` - the complete server config.
- `model-sha256.txt` - model, mmproj, MTP and expert-profile hashes.
- `bench_report.py` - the measurement script (runs unchanged against any Strata server).

Key numbers (warm, single-sequence, 256-token cap, greedy, reasoning off): prompt 1,436.9 tok/s median at 17,971 tokens; 1,446.4 at 44,011; 1,410.5 single run at 119,731. Decode ~73 tok/s median. Two configuration notes measured for this rig: `--prefill auto:32768` (the default 8192 chunk cap costs 50-100% prompt speed on helper-cache rigs - each >=1024-token chunk streams every non-resident expert) and `STRATA_PREFILL_CPU_SHARE=auto` (-23% TTFT on sub-1024-token prompts, output-identical in a fixed-prompt check).

No engine changes in this PR - results only.

## Extra Notes

The rig has no P2P between the cards (different CPU sockets); decode includes the helper round-trip. The 0.1.39-to-0.1.40 decode delta and the helper-cache A/B were reported earlier in #1079.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。