Pull requests / #1695

#1695 bench: P4 prefill 78.8 -> 114.8 tok/s with existing auto sizing; expert reuse roofline

open · @CC-David-CC · 0 comentários · No GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Descrição

Independent P4 prefill measurements inspired by @vizdatom's #1670. Results-only: this does not implement or run the proposed layer-major path.

| Prompt | Fixed 512, median of 3 | Existing auto, one run |
| --- | ---: | ---: |
| 4,096 | 78.018453 tok/s | 107.259667 tok/s |
| 8,192 | 78.840339 tok/s | 114.753823 tok/s |

One Tesla P4, Flash-Next-GSQ-RCO IQ3_XXS, int8 KV, CUDA 12.0, dual Xeon E5-2697 v3. Base fb58e0d. Warm page cache, zero prompt reuse, one generated token/request. Auto selected up to 1,792 tokens. Instrumented controls disable CPU expert sharing; matching fixed-512 controls were 77.928650/78.933397 tok/s.

8K fixed-512 copied 262.788941 GB of expert weights, 30.871322 GB distinct (88.25% repeated payload). Auto copied 193.275469 GB. Measured pinned H2D bandwidth: 11.886827 GB/s; VRAM read: 147.402334 GB/s. Copy waits were 6.2% of the fixed-512 timeline; fewer bytes are not an equivalent wall-time gain.

Includes exact prompts/output IDs, all 16 request records, configs, logs, replay driver, bandwidth sources, analysis and chart. The opt-in 21-line counter patch is a reproduction artifact only; engine source is unchanged. CUDA tested; no HIP/SYCL build or multi-token quality gate. Model revision/full asset hashes were not recorded.

Full report: bench/results/2026-10-09-community-p4-prefill-roofline/README.md

No site

Links install, modelos, releases.