Pull requests / #850

#850 bench: community report, Tesla V100 32 GB + RTX 4070 layer split with 16 GB of RAM (IQ2_XS)

closed · @christopherrobertbrooks-tech · 0 comentarios · En GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants

Descripción

Community benchmark report: the full Qwen3.8-Flash-Next **IQ2_XS** split across a **Tesla V100 32 GB (sm_70, PCIe 3.0 x4) and an
RTX 4070 12 GB (sm_89)**, in both card orders, on a PC with **16 GB of RAM**. In `bench/results/2026-10-04-community-v100-rtx4070-split/`.
Results only, no code changes (companion to #823, the V100 alone).

- **Build:** engine 0.1.39 from source with `--cuda 12`, `"archs": [70, 89]`.
- **Configuration:** `"layer_split": "auto"`, low-RAM mode (`--mmap-experts`), context 131,072, KV int8, MTP `--spec 4`, no vision.
  Auto chose K=39 with the V100 first and K=11 with the 4070 first.
- **Method:** the 2x MI50 report's `benchmark.py`, unchanged (3 runs each at 4,096 / 32,768 / 128,000 fresh prompt tokens,
  256-token cap, reasoning none) + needles 32k/128k x 10/50/90.
- **Results (medians), V100 first:** decode 84.9 / 81.5 / 75.1 tok/s, prompt 996 / 1,602 / 1,552 tok/s. 4070 first: decode
  79.1 / 83.8 / 77.1, prompt 985 / 1,535 / 1,533. Needles 6/6 in both orders.
- **Mixed-arch split on long prompts:** every 32K/128K prompt went through the second card in both orders, no errors (the
  failure in #690 is on a mixed AMD pair).
- **PCIe:** the V100's x4 link peaks at ~3.6 GB/s (essentially full) during long-prompt reading.
- **Limitations:** one machine, one size, three runs per configuration, no vision in the speed runs.

The report was drafted with an AI assistant (Claude) from our measurements, and checked by me.

En el sitio

Enlaces a install, modelos, releases.