Pull requests / #839

#839 Add experimental Linux gfx900 setup with wave64 validation

closed · @sagentlab · 0 comentarios · En GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Descripción

Vega 10 (`gfx900`) is rejected by setup, and the existing wave64 compatibility backend needs fixes to compile with the current HIP toolchain. This adds an explicit Linux opt-in that installs a gfx900-capable ROCm SDK, builds the wave64 engine, checks the device before downloading the model, and runs Qwen3.8-Flash-Next IQ2_XS on a 16 GB Vega 10 PC.

Use `STRATA_EXPERIMENTAL_GFX900=1 ./setup.sh --yes --backend hip`. Ordinary setup retains its existing supported architectures and ROCm pin. The experimental build uses a separate `build-gfx900` directory and AMD's TheRock gfx900 packages pinned to `7.14.0a20260612`; Windows and mixed wave32/gfx900 builds are refused. Installation, reproduction commands and measured limits are documented in `docs/AMD_GFX900.md`.

The wave64 changes fix HIP/libstdc++ attribute conflicts and kernel-template arguments in `cudaFuncSetAttribute`, preserve modulo-2^32 signed software dp4a accumulation, and check the compiled architecture and wave size before launching kernels. New GPU checks exercise both 32-lane halves of a wave64 and compare BF16/FP16 GEMM against CPU references.

Validation on 2026-10-04: gfx900, PCI 1002:6860, 16 GB HBM, Xeon W-2191B, 125 GiB RAM, Arch Linux and the existing amdgpu driver:

- Release engine build, device allocation self-test, wave64 intrinsic check and BF16/FP16 GEMM checks pass.
- 94 setup tests pass across gfx900, choices, config, older GPU, dependency pin and update suites; all 26 ROCm SDK installation checks pass with the installer environment.
- Engine CTest: 40/42 pass. `ple_parity` requires an unavailable external Q2_0 fixture; `platform_memory_test` requests a 256 MiB lock above this account's 8 MiB limit. Neither is counted as passing. Numerical expert, attention, routing, sampler and KV checks pass.
- The existing AMD setup suite passes 27/28 on Linux. Its Windows prebuilt fixture writes `strata.exe` while the host expects `strata`; the same failure was reproduced with unchanged upstream setup.py.
- Complete installer succeeds: both model shards match pinned SHA-256 hashes, the MTP head packs, and the server loads with context 8192, reserve 3072 MiB, no vision, 17 AVX2 expert workers and an 8.23 GiB GPU expert cache.
- Live HTTP checks pass for health, model discovery, browser UI, arithmetic and Python code, repeated-prompt reuse, streaming completion, Anthropic Messages, OpenAI Responses and monitoring. A further 256-token binary-search explanation completes successfully.

This is an experimental port validated on one PC. Images and multiple-card inference remain unvalidated; these checks do not establish compatibility or throughput for every Vega 10 board, model or context size.

En el sitio

Enlaces a install, modelos, releases.