Pull requests / #839

#839 Add experimental Linux gfx900 setup with wave64 validation

closed · @sagentlab · 0 comments · View on GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

Vega 10 (`gfx900`) is rejected by setup, and the existing wave64 compatibility backend needs fixes to compile with the current HIP toolchain. This adds an explicit Linux opt-in that installs a gfx900-capable ROCm SDK, builds the wave64 engine, checks the device before downloading the model, and runs Qwen3.8-Flash-Next IQ2_XS on a 16 GB Vega 10 PC.

Use `STRATA_EXPERIMENTAL_GFX900=1 ./setup.sh --yes --backend hip`. Ordinary setup retains its existing supported architectures and ROCm pin. The experimental build uses a separate `build-gfx900` directory and AMD's TheRock gfx900 packages pinned to `7.14.0a20260612`; Windows and mixed wave32/gfx900 builds are refused. Installation, reproduction commands and measured limits are documented in `docs/AMD_GFX900.md`.

The wave64 changes fix HIP/libstdc++ attribute conflicts and kernel-template arguments in `cudaFuncSetAttribute`, preserve modulo-2^32 signed software dp4a accumulation, and check the compiled architecture and wave size before launching kernels. New GPU checks exercise both 32-lane halves of a wave64 and compare BF16/FP16 GEMM against CPU references.

Validation on 2026-10-04: gfx900, PCI 1002:6860, 16 GB HBM, Xeon W-2191B, 125 GiB RAM, Arch Linux and the existing amdgpu driver:

- Release engine build, device allocation self-test, wave64 intrinsic check and BF16/FP16 GEMM checks pass.
- 94 setup tests pass across gfx900, choices, config, older GPU, dependency pin and update suites; all 26 ROCm SDK installation checks pass with the installer environment.
- Engine CTest: 40/42 pass. `ple_parity` requires an unavailable external Q2_0 fixture; `platform_memory_test` requests a 256 MiB lock above this account's 8 MiB limit. Neither is counted as passing. Numerical expert, attention, routing, sampler and KV checks pass.
- The existing AMD setup suite passes 27/28 on Linux. Its Windows prebuilt fixture writes `strata.exe` while the host expects `strata`; the same failure was reproduced with unchanged upstream setup.py.
- Complete installer succeeds: both model shards match pinned SHA-256 hashes, the MTP head packs, and the server loads with context 8192, reserve 3072 MiB, no vision, 17 AVX2 expert workers and an 8.23 GiB GPU expert cache.
- Live HTTP checks pass for health, model discovery, browser UI, arithmetic and Python code, repeated-prompt reuse, streaming completion, Anthropic Messages, OpenAI Responses and monitoring. A further 256-token binary-search explanation completes successfully.

This is an experimental port validated on one PC. Images and multiple-card inference remain unvalidated; these checks do not establish compatibility or throughput for every Vega 10 board, model or context size.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.