Pull requests / #121
#121 Add gfx1100 HIP support with tuned prefill and bounded expert uploads
closed · @2jztricks · 0 commentaires · Sur GitHub
BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
RX 7900 XTX users need a native HIP path and bounded host-memory transfers to serve Flash Next reliably. This replaces #94 with the complete opt-in gfx1100 backend rebased onto current main, adds tuned prefill, and includes fresh performance evidence. - Retains HIP runtime/wave32 compatibility, native mmap layout validation, MTP and numerical handoff checks. - Enables pinned-GGML MMQ on HIP and guarded hipBLASLt tuning for gfx1100/version 100100. Adds opt-in native embedding/QSA batching, selectively adapted from #108 (acd4872), preserving upstream borrowing and CUDA paths. - Adds static CPU-resident expert complements with available-RAM/cgroup guards, a POSIX PLE read pool, and reusable pinned HIP expert-upload staging. The staging fix is used for both startup and borrowed-cache refills; an intermediate package without it stalled and was rejected. Fresh 2026-09-29 measurements on RX 7900 XTX at the unchanged 272 W cap, Orca Flash Next IQ3_XXS, MTP 4, configured 262144 context, 8192 prefill batch: | Request | Prefill tokens/s | Output tokens/s | Total request time | | --- | ---: | ---: | ---: | | 4210 tokens, first use | 447.5 | 55.5 | 11.729 s | | 8830 tokens, first use at this size | 873.7 | 55.7 | 12.420 s | | 4210 tokens, warmed | 901.1 | 56.7 | 6.937 s | | 8830 tokens, warmed | 966.5 | 59.2 | 11.314 s | Each fresh prompt reused zero KV tokens and generated the intentional 128-token cap. All four fresh prompts and four cached follow-ups completed. A newly measured same-session existing HIP runtime control gives 1.80–2.02x prefill and 40–46% lower fresh-request wall time. This is not a comparison against unmodified upstream, and configured 262K capacity is not a full-context benchmark. Validation: HIP and CUDA executable builds; 29/29 selected HIP CTests; three real IQ3_XXS PLE matrix graph replays; POSIX direct-file queued reads/short EOF/wake/close/reopen test. The two excluded CTests and all measurement conditions, exact source/binary identifiers and sanitized raw timings are documented in `docs/AMD_HIP_PERFORMANCE.md`. CUDA has compile validation here, not a new CUDA throughput run. These numerical and capped-throughput tests do not establish coding quality, full-context/vision behavior or mixed-vendor execution. Supersedes #94. Related: #48. No model weights or personal serving/gateway configuration are included.
Sur le site
Liens install, modèles, releases.