Pull requests / #628

#628 bench: report Coder IQ1_M on Threadripper 3990X and RX 6900 XT

closed · @Yasei-no-otoko · 0 コメント · GitHub で見る

BenchmarksAMD / HIPModels & quantsWindows

本文

Add reproducible Coder IQ1_M measurements from a Ryzen Threadripper 3990X (64 cores / 128 threads, 128 GiB DDR4-3200) and Radeon RX 6900 XT (16 GiB, gfx1030) on Windows 11 Pro 10.0.26300, AMD driver 32.0.21045.5002.

The report separates three sets of evidence:

- Official v0.1.38 release-binary sweep: 63/31/15 workers, including three-run short, 4K and 32K workloads at 63/31 workers. The historical manual 31-worker selection is retained as a measurement, not a proposed default.
- Matched-setting source builds before/after the processor-group correction in #626: with 63 explicit workers and 3,245 effective cache slots, short/4K decode medians improved from **12.69 / 13.08** to **46.09 / 47.43 tokens/s** (about **3.6x**). All 31-worker comparison rows are included, including the fixed-31 regression. A separate single 32K request measured 48.87 tokens/s.
- Pre-review code with automatic worker selection: **63 workers**, automatic cache at 2,675 slots, short/4K medians **43.13 / 57.91 tokens/s**. This has a different cache budget and is explicitly separated from the before/after comparison.

Each short/4K case uses three fresh requests, 256 output tokens, temperature 0 and no reasoning; warm-up is excluded and all measured prefix reuse is zero. Context is 65,536 with INT8 KV and 32,768 resident cells. The data includes per-request input/output, engine and client timings, median/min/max summaries, configs, source/binary/model hashes, topology and hardware snapshots, build/test records, and standard-library Python reproduction scripts. The separate generated clamp smoke test passed 2,001 bounded integer cases.

The machine reports one NUMA node spanning two processor groups. This is one interactive desktop and synthetic coding workload: background applications stayed open, free VRAM varied and the fixed-cache patched runs reported low VRAM. Results do not establish an optimum for all hardware or general coding quality. The final automatic run and single 32K observation are not pooled into the paired medians.

This PR contains benchmark data/documentation only; the engine change and tests are in #626. Neither PR changes the automatic default to 31 workers.


Review follow-up: local account names in the captured CPU-test paths are redacted. The added review-validation record covers the later reversible host CPU Set implementation in #626; historical throughput figures remain tied to their recorded revisions and were not rerun for that follow-up. Both AVX-512 pool tests, including pool_stress, skip their workloads on this CPU; its zero exit status was previously mistaken for a stress-workload pass, and the report now corrects that distinction.

関連リンク

インストール・モデル・リリースへの站内リンク。