Pull requests / #850
#850 bench: community report, Tesla V100 32 GB + RTX 4070 layer split with 16 GB of RAM (IQ2_XS)
closed · @christopherrobertbrooks-tech · 0 评论 · 在 GitHub 查看
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
描述
Community benchmark report: the full Qwen3.8-Flash-Next **IQ2_XS** split across a **Tesla V100 32 GB (sm_70, PCIe 3.0 x4) and an RTX 4070 12 GB (sm_89)**, in both card orders, on a PC with **16 GB of RAM**. In `bench/results/2026-10-04-community-v100-rtx4070-split/`. Results only, no code changes (companion to #823, the V100 alone). - **Build:** engine 0.1.39 from source with `--cuda 12`, `"archs": [70, 89]`. - **Configuration:** `"layer_split": "auto"`, low-RAM mode (`--mmap-experts`), context 131,072, KV int8, MTP `--spec 4`, no vision. Auto chose K=39 with the V100 first and K=11 with the 4070 first. - **Method:** the 2x MI50 report's `benchmark.py`, unchanged (3 runs each at 4,096 / 32,768 / 128,000 fresh prompt tokens, 256-token cap, reasoning none) + needles 32k/128k x 10/50/90. - **Results (medians), V100 first:** decode 84.9 / 81.5 / 75.1 tok/s, prompt 996 / 1,602 / 1,552 tok/s. 4070 first: decode 79.1 / 83.8 / 77.1, prompt 985 / 1,535 / 1,533. Needles 6/6 in both orders. - **Mixed-arch split on long prompts:** every 32K/128K prompt went through the second card in both orders, no errors (the failure in #690 is on a mixed AMD pair). - **PCIe:** the V100's x4 link peaks at ~3.6 GB/s (essentially full) during long-prompt reading. - **Limitations:** one machine, one size, three runs per configuration, no vision in the speed runs. The report was drafted with an AI assistant (Claude) from our measurements, and checked by me.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。