Pull requests / #927

#927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854

closed · @xjc10 · 0 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

A community benchmark report, data only: one new folder under `bench/results/` and its line in `docs/COMMUNITY_BENCHMARKS.md`. No engine or setup changes.

**Machine:** 2x AMD Radeon RX 6900 XT 16 GB (gfx1030, PCIe 4.0 x8/x8 by the measured 14.1 GB/s per card), Ryzen 5 5600X (AVX2), 128 GB DDR4, Ubuntu 26.04, ROCm 10.0.0. **Model:** the original Flash-Next GSQ-RCO IQ3_S (setup's files, hashes checked), 131,072-token context.

**What was run:** the RTX 5090 report's `benchmark.py` unchanged (three runs each at 4,096 / 32,768 / 128,000 prompt tokens, 256-token greedy outputs, zero reused tokens) and `tools/needle_bench.py` (six needles), in seven configurations: one card, the layer split across both cards and the expert-helper mode (`--expert-cache-device1 auto --remote-expert-opt`), each on stock `6f32ec0` and with #835, #849 and #854 merged, plus the helper mode's `--pcie-frac 0` workaround. Every configuration has its engine log, results, summary, needles, one-second telemetry (an AMD sysfs form of the report's `monitor.py`), status, config and build record in the folder.

**Medians:**

| | prompt tok/s at 4K / 32K / 128K | decode tok/s | TTFT at 128K |
| --- | ---: | ---: | ---: |
| one card, stock | 457 / 474 / 459 | 40-43 | 279 s |
| one card, with the PRs | 868 / 1,100 / 1,023 | 40-43 | 125 s |
| layer split, stock | 446 / 723 / 821 | 58-61 | 156 s |
| layer split, with the PRs | 853 / 1,616 / 1,823 | 57-59 | 70 s |
| helper, stock (probed share 0.39) | 457 / 473 / 458 | 38-42 | 279 s |
| helper, stock, `--pcie-frac 0` | 460 / 473 / 458 | 56-63 | 280 s |
| helper, with the PRs (probed share, #854) | 875 / 1,101 / 1,024 | 60-65 | 125 s |

All 42 needles were found. No request failed or stalled.

**Two things the README states up front:** every configuration ran with `--adapt-every 100000`, because the adaptive swaps hang the engine on this machine after a long prompt (#884) - it costs a few percent of decode; and the PR numbers are for pull requests that were open when this was measured (#835 changes the prompt path's arithmetic, #849 and #854 do not change any value).

Developed with an AI coding assistant; every number above was measured on 2x RX 6900 XT (gfx1030) / Ryzen 5 5600X, ROCm 10.0.

Mehr auf der Site

Links zu Install, Modellen, Releases.