Pull requests / #858
#858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
closed · @agurrrrr · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationLinux
本文
## What this fixes
`STRATA_HIP_GFX906=ON` does not build at `main` (`6f32ec0`, v0.1.39). The fused_gr v3 read that `cfd3b72` put on
main calls `cudaFuncSetAttribute` on kernel templates with more than one argument, and sets the shared-memory
carveout: (`cfd3b72` is the same work as PR #646 - same author, same subject line - but it went onto main as a direct
commit; PR #646 itself was closed unmerged on 2026-10-04, so the calls below are main's own, from `cfd3b72`.)
```cpp
cudaFuncSetAttribute(gr_down_v3_kernel<1, kFusedGrMaxT, false>, cudaFuncAttributeMaxDynamicSharedMemorySize, need1)
cudaFuncSetAttribute(gr_down_v3_kernel<1, 4, true>, cudaFuncAttributePreferredSharedMemoryCarveout, 100);
```
f166564 added the carveout name to the wave32 compat header (`include/strata/hip_compat/cuda_runtime.h`), which
already has `cudaFuncSetAttribute` as a template function. The gfx906 compat header
(`include/strata/platform/hip_compat/strata_hip.h`) has neither:
- `cudaFuncSetAttribute` is a function-like macro there, so the preprocessor splits `gr_down_v3_kernel<1, 4, true>`
at its commas into extra macro arguments;
- `cudaFuncAttributePreferredSharedMemoryCarveout` is not mapped.
The build stops at fused_gr.cu with 12 errors, the same ones for each of the five calls on lines 1089-1093:
```
src/kernels/cuda/fused_gr.cu:1089:81: error: too many arguments provided to function-like macro invocation
src/kernels/cuda/fused_gr.cu:1089:17: error: use of undeclared identifier 'cudaFuncSetAttribute'; did you mean 'hipFuncSetAttribute'?
src/kernels/cuda/fused_gr.cu:1089:133: error: comparison between pointer and integer ('hipError_t (*)(const void *, hipFuncAttribute, int)' ...) and 'hipError_t')
...
19 warnings and 12 errors generated when compiling for gfx906.
```
The change (1 file, +5 / -1): the macro becomes a template function of the same shape as the wave32 header's, and the
carveout name maps to `hipFuncAttributePreferredSharedMemoryCarveout`. The header is on the include path only under
`STRATA_HIP_GFX906` (the `STRATA_HIP_GFX906` block is `CMakeLists.txt:106-124`, and the three include-path entries are
116, 119 and 120), so the CUDA build and the wave32 HIP build do not see it.
## How it was checked
AMD Instinct MI50 32 GB (gfx906, PCI id 1002:66a1; `rocm-smi` reports 34,342,961,152 B of VRAM), EPYC 7452,
125 GiB RAM; CachyOS (kernel 7.2.8), ROCm 7.2.4 (HIP 7.2.53211,
AMD clang 22.0.0git).
| | result |
|---|---|
| `main` (`6f32ec0`), `ninja .../fused_gr.cu.o` | 12 errors (above) |
| this branch, `ninja` (all targets, `STRATA_BUILD_TESTS=ON`) | 269 / 269, 0 errors, 32 s |
```sh
G=--gcc-install-dir=/usr/lib/gcc/x86_64-pc-linux-gnu/14.3.1
cmake -S . -B build-gfx906 -G Ninja -DSTRATA_HIP_GFX906=ON -DCMAKE_HIP_ARCHITECTURES=gfx906 \
"-DCMAKE_HIP_FLAGS=--rocm-path=/opt/rocm $G" "-DCMAKE_CXX_FLAGS=$G" -DCMAKE_HIP_COMPILER_ROCM_ROOT=/opt/rocm \
-DSTRATA_GGML_DIR=<llama.cpp at 3cf0325> \
-DCMAKE_C_COMPILER=/opt/rocm/llvm/bin/clang -DCMAKE_CXX_COMPILER=/opt/rocm/llvm/bin/clang++ \
-DCMAKE_BUILD_TYPE=Release -DSTRATA_BUILD_TESTS=ON
ninja -C build-gfx906
```
The three flags beyond `docs/AMD_HIP.md`'s command are for this distro's ROCm packaging, not for this change: clang
sits in `/opt/rocm/lib/llvm`, so CMake guesses the ROCm root as `/opt/rocm/lib`, fails to identify the HIP compiler
and drops `-O3 -DNDEBUG` (hence `CMAKE_HIP_COMPILER_ROCM_ROOT` and `--rocm-path`); and GCC 16's `<format>` spells
`[[__gnu__::__noinline__]]`, which HIP's `__noinline__` macro breaks (hence GCC 14's headers).
**On the card.** An engine built from the same source tree (byte-identical to this branch) has served on the MI50
since 2026-10-05: Qwen3.8-Flash-Next GSQ-RCO IQ3_S, `--max-context 262144 --kv int8 --spec 4`, the experts in host
RAM. Every other `cudaFuncSetAttribute` call in the gfx906 build (`fused_gr.cu:725-731`, `qsa.cu`, the prefill
kernels) now goes through the template and makes the same `hipFuncSetAttribute` call as the macro did. Measured on that
server, greedy (a range is two runs, a single number one run):
| | tok/s |
|---|---:|
| decode, short prompts (Korean prose / code) | 33.9-34.6 / 39.8-40.9 |
| prompt, 4,526 tokens | 375-378 |
| prompt, 42,705 tokens | 393 |
| prompt, 78,284 tokens | 387 |
A needle in the 43K and in the 78K prompt is found, and an OpenAI tool call returns the right arguments.
Where those numbers come from: the short-prompt and the 4,526-token rows are the `iq3s-0139-alone-mi50` entries of
my bench log (`~/code/local-llm/bench/strata-vs-golbang/results.jsonl`; greedy, 768 generated tokens, two reports per
prompt), and the 42,705 and 78,284-token rows are long-prompt needle runs against the same server on the same day.
## Not covered
- `STRATA_GR_V3=1`, the opt-in read that sets the carveout, was not run on the card; this PR only lets it compile.
- `ctest` was not run on the card for this PR: the card serves the model above with 586 MiB of VRAM left.
- The CUDA and the wave32 HIP builds were not rebuilt: neither includes this header.
関連リンク
インストール・モデル・リリースへの站内リンク。