Pull requests / #1638
#1638 bench: add single R9700 performance report
open · @zihaomu · 0 Kommentare · Auf GitHub
BenchmarksAMD / HIPModels & quantsDocumentationLinux
Beschreibung
## Summary Publish a single-R9700 Linux performance report for original Qwen3.8-Flash-Next IQ3_S, with IQ2_XS as a separate contrast. Measured on 2026-10-08 with a frozen experimental 0.1.40.3 build, ROCm 7.2.3, one 32 GB R9700, two EPYC 9334 CPUs and about 503 GiB RAM. Five controlled IQ3_S pairs for #1107 reduce native TTFT by 16.18%/14.07% at +512/+900 tokens after 32K history. The report preserves the +256 timing spread and fresh/decode regressions; it does not establish an all-workload improvement. ## Measured single-GPU speed **GPU: 1 × AMD Radeon AI PRO R9700, 32 GB VRAM (single GPU), gfx1201.** Single-card isolation was verified using ROCr UUID isolation and HIP device-count/BDF checks. The approximately 503 GiB stated above is host system RAM, separate from GPU VRAM. Original Flash-Next IQ3_S on the R9700/EPYC host above, using auto expert cache and prefill, MTP, int8 KV and Tensile. Concurrency 1, greedy decoding, 256 generated tokens per request. Each input size has three measured repeats after warmup; every input is fully processed with no prefix reuse. Brackets show the full minimum–maximum range. | Input tokens | Prompt tok/s median [range] | Decode tok/s median [range] | Native TTFT seconds median [range] | Total seconds median | | ---: | ---: | ---: | ---: | ---: | | 1K (1,024) | 632.9 [514.8–634.0] | 101.9 [101.2–102.8] | 1.64 [1.64–2.01] | 4.15 | | 4K (4,096) | 768.1 [716.3–768.8] | 96.4 [94.4–100.6] | 5.35 [5.35–5.74] | 8.05 | | 32K (32,768) | 750.5 [750.4–775.4] | 95.6 [90.4–98.2] | 43.69 [42.29–43.70] | 46.29 | | 128K (131,072) | 718.4 [713.9–719.0] | 87.6 [85.3–91.4] | 182.48 [182.33–183.64] | 185.39 | Native TTFT measures request submission to the first engine token. Timings exclude model loading and warmup; TTFT excludes HTTP/SSE and client networking. Filesystem and expert caches are warmed. These are absolute speeds of the frozen experimental build, separate from the controlled #1107 comparison. [Full report and per-run data](https://github.com/zihaomu/Strata/blob/5562193c68fc5c88bff0710ebd28823b710a4298/bench/results/2026-10-08-community-r9700-linux/README.md#fresh-iq3_s-speed) ## What changed - Add a focused report, short Chinese summary, reproduction commands and community-index link. - Keep readable per-request and summary CSVs plus one lossless measurement bundle: all 94 paired + 12 absolute-speed formal requests, all recorded warmups (213 records total), output IDs/text, session metadata and 23 complete engine logs. - Keep only the required performance scripts, exact token fixtures, configurations, asset hashes and archival patches that reconstruct the measured source from public release commit `d5ea713`. Active engine/backend code and runtime defaults are unchanged. - Link historical tuning, numerical diagnostics and optional HTTP checks to the [original full archive](https://github.com/zihaomu/Strata/tree/17006083063b4443458f6f0b3f5c8325cf9cca2a/bench/results/2026-10-08-community-r9700-linux), outside the final report diff. ## Extra Notes Validation: `verify_report.py --check-sources` passed (29 published-file hashes, source reconstruction, all formal requests and CSV/statistics checks). The 61 retained JSON objects and 23 decompressed logs match the original archive. Extraction round-trip, detection of omitted CSV rows/changed output IDs, documented shell/Python syntax, local links and `git diff --check` passed. Measurements are of the frozen experimental source, not newer upstream main. No GPU inference was rerun while preparing this report. Absolute speeds and controlled gains use different settings. Hardware dependence, timing boundaries and missing coverage are stated in the report. Related engine evidence stays in #1107, #1474 and #1540; the archived architecture-scoped signed-zero fix is not a test of #1540's broader patch.
Mehr auf der Site
Links zu Install, Modellen, Releases.