Pull requests / #1278
#1278 HIP: run the K-quant MMQ test on HIP builds; UD-Q4_K_XL measured on an R9700
open · @ConnorHaggerty · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
描述
Description:
## Summary
- **Test wiring:** register `tests/cuda/prefill_mmq_kquant_test.cpp` in the HIP MMQ branch as `hip_prefill_mmq_kquant` (named and linked like
`hip_prefill_fused_iq`). Since 43881f1 the HIP build compiles the Q4_K / Q5_K / Q5_1 / Q6_K MMQ instances under `-DSTRATA_MMQ_KQUANTS=ON`, but the test
that checks those products was only registered for CUDA.
- **docs/UNSLOTH_Q4.md:** the setup bullet said UD-Q4_K_XL had not run on AMD because its prompt kernels were NVIDIA-only, and that images were not wired
to it. Neither holds on 0.1.40. Reworded it, and added an "On an AMD card (R9700, set up by hand)" section with what was measured and what was not. Setup's
behaviour (`nvidia_only`, `"vision": False`) is unchanged: setup itself has not been run with this file on AMD.
## Tested on
Radeon AI PRO R9700 (32 GB, gfx1201), i7-12700KF (AVX-2, no AVX-512), 96 GB DDR4-3200, Ubuntu 26.04, setup's ROCm 7.10 wheels, `main` at 82f46a8 plus this
commit, built with `-DSTRATA_MMQ_KQUANTS=ON -DSTRATA_BUILD_TESTS=ON`.
- `hip_prefill_mmq_kquant` passes. Max rel_l2: Q4_K 0.0030, Q5_K 0.0034, Q6_K 0.0016 (gate/up); Q5_1 0.0015, Q8_0 0.0005 (down).
- Full ctest: 83 of 89 pass or skip. The 6 failures don't link anything this PR touches:
- `expert_cache_segmented_test`: `--vram-elastic` is CUDA-only.
- `ple_parity`: needs a Q2_0 file that isn't on this PC.
- `platform_memory_test`: `ulimit -l` is 8 MiB here.
- `expert_multi_test`: needs AVX-512.
- `hip_q2_zero`: 128 negative zeros come out as +0.
- `hip_prefill_wmma_gemm_parity`: the WMMA kernels are gfx11-only, so on gfx1201 every case is declined and the test reports FAILED ("zero test cases
passed") instead of skipping. Probably wants exit 77 off gfx11, like `hip_prefill_fused_moe`.
- UD-Q4_K_XL at revision 38bb39e (all four sha256 checked), `--resident-budget-gib 64`, 262K context, `--kv int8 --kv-resident 32768`, MTP draft layer:
- prompts of 19K-63K tokens: 800-941 tok/s
- one 1,000-token answer: 32-37 tok/s
- two answers at once: ~11 tok/s each
- images on the CPU encoder at 1,024 tokens: correct answers about a screenshot's small text
- no watchdog stalls in 12 cold long-prompt runs, 6 of them with a second request arriving mid-prompt on 2 slots. Engine 0.1.39 stalled in about half of
such runs on this PC with the RAM complement mapped for the GPU; 0.1.40 did not.
- Not done: the greedy comparison against llama.cpp on AMD (the doc says so).站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。