Issues / #881

#881 [Windows / AMD] RX 6800M (gfx1031) runs with a self-built HIP engine — plus 2 Windows bugs found (packaging encoding, vision `vcvars`)

closed · @Hugua700 · 3 评论 · 在 GitHub 查看

Setup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

描述

Following `docs/AMD_HIP.md` ("the ready-made zip itself has not run a model on a discrete card yet — please report" / "Reporting a Windows AMD run"), here is one — on a card the current prebuilt engine does **not** cover.

## Hardware / software

| | |
|---|---|
| GPU | AMD Radeon RX 6800M (Navi 22, **gfx1031**, 12 GB, 20 CU); integrated Radeon (gfx1036) idle |
| Driver | AMD Adrenalin `32.0.21045.5002` (Windows reports PCI `1002:73DF`) |
| Host | Windows 11 Home, Ryzen 9 5900HX (AVX2), 31.4 GB RAM |
| Strata | source `0.1.39`, engine `0.1.39`, ROCm `10.2.0a20260930` (TheRock) |
| Model | Qwen3.8-Flash-Next **Coder IQ1_M**, `--max-context 32768`, `--kv int8`, `--resident-experts` (low-RAM mode) |

## 1. `setup.py` refuses the card, for two independent reasons

`START-HERE.bat --backend hip` prints:

```
GPU 1: AMD Radeon RX 6800M, 12 GB VRAM - not supported ... this is unknown (PCI 73DF)
```

1. `_WIN_AMD_DID` has no `0x73DF` entry (only `0x73BF/0x73AF/0x73A5 → gfx1030`).
2. `_WIN_AMD_NAME`'s fallback **deliberately excludes** the mobile part: `RX\s*6800(?!\s*[MS])`.

Reason (2) is correct in spirit — the mobile **6800M is Navi 22 = gfx1031**, while desktop 6800/6900 are Navi 21 = gfx1030. But it may bite users: the README's "RX 6800 / 6900 series" reads as covering the 6800M, and the prebuilt Windows zip does not contain gfx1031 either, so the card is refused twice over:

```
BUILD.json  archs:           [gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030]
rocm/.kpack/ blas_lib_gfx*.kpack:      gfx1030, gfx1100, gfx1101, gfx1102, gfx1200, gfx1201
```

## 2. What worked: self-build for gfx1031

`rocm-sdk-device-gfx1031` exists in the TheRock index, so I mapped `0x73DF → gfx1031` in `_WIN_AMD_DID` and built the engine myself (`tools/hip/build_windows.bat` with `STRATA_HIP_ARCHS=gfx1031`) — no other source change needed for the engine itself. Then `START-HERE.bat --backend hip --prebuilt dist\`.

`strata-device.exe --selftest`:

```
device 0: AMD Radeon RX 6800M
  HIP arch            gfx1031 wave32 (compiled for gfx1031)
  multiprocessors     20
  VRAM total / free   11.984 GiB / 11.068 GiB
  driver / runtime    71726391 / 71726391
  HIP runtime         <engine>\amdhip64_7.dll
  plan + KV vs free   FITS
strata-device selftest OK
```

Running (Coder IQ1_M, low-RAM mode, GPU holds ~30 % of the experts):

```
/health → {"status":"ok","loaded":true,"max_context":32768,"images":true}
short English/code answer   ~8.5 s
short Chinese answer        ~9.1 s
tool calling                finish_reason="tool_calls"; correct follow-up answer using the tool result
image (720x360, big digits) correct read, ~20 s including CPU encoding
```

As expected for RDNA2, setup reports no hipBLASLt table → plain hipBLAS for prompt GEMMs.

## 3. Bug: `tools/hip/package_windows.py` fails on a non-UTF-8 locale (compile succeeds, packaging always fails)

Line ~145/146 call `Path.read_text()` without an encoding, so the locale codec is used. On a Chinese Windows (GBK) `CMakeLists.txt` cannot be decoded:

```
[131/131] Linking CXX executable strata.exe          <-- build is fine
Traceback (most recent call last):
  File "...\tools\hip\package_windows.py", line 145, in main
    version = re.search(r"project\(strata VERSION ([\d.]+)", (ROOT / "CMakeLists.txt").read_text()).group(1)
UnicodeDecodeError: 'gbk' codec can't decode byte 0xbf in position 2: illegal multibyte sequence
```

⇒ the engine compiles and then the zip is never produced. Fix: pass `encoding="utf-8"` (2 call sites here; `setup.py` has ~31 more unprotected `read_text()` calls that are exposed to the same class of failure — running setup with `PYTHONUTF8=1` works around all of them).

## 4. Bug: the CPU vision encoder can never be built on Windows (`vcvars=None`)

`build_vision_cpu()` calls:

```python
cmake_build(ROOT / "tools" / "vision", ROOT / "build-vision", "strata-vision",
            [f"-DLLAMA_DIR={llama}", "-DSTRATA_VISION_CUDA=OFF"], None, "")
```

On Windows `cmake_build()` writes a batch file whose first line is `call "{vcvars}"` — with `vcvars=None` that is `call "None"`, so the build always fails. The NVIDIA path passes a real `vcvars` (`setup.py:2323`), so this only shows up if the HIP path ever reaches it.

It currently cannot: the "engine has no encoder → compile it" branch is guarded by `if eng is not None and not hip`, so the **Windows-HIP prebuilt path never builds or installs an encoder at all** — even if `hip_vision()` were changed to allow `--vision cpu`.

## 5. …but the CPU image encoder *does* work on Windows

`hip_vision()` refuses `--vision cpu` on Windows with "the CPU image encoder is Linux-only for now (… has no image encoder)". As far as I can tell that is a packaging/plumbing gap rather than a capability one:

* `tools/vision` is a standalone CMake project (`add_executable(strata-vision strata_vision.cpp)`, links llama.cpp's `mtmd` + `llama`), **not** a HIP/CUDA target;
* its `CMakeLists.txt` has an explicit MSVC branch (`if(STRATA_PORTABLE AND MSVC)`), and `strata_vision.cpp` contains `#ifdef _WIN32`;
* the prebuilt Windows HIP zip simply doesn't contain `strata-vision.exe` (`PROGRAMS = ("strata.exe", "strata-device.exe")`).

I built it with MSVC + Ninja (`-DSTRATA_VISION_CUDA=OFF -DSTRATA_PORTABLE=ON`), dropped it into `engine\`, recorded `vision: "cpu"` + `vision_src` in `engine/BUILD.json`, downloaded the `mmproj`, and added the config section. It starts and reports:

```
READY 2560
strata-vision: on the CPU, 8 threads, no warm-up
load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
```

Then `--vision` + the config's `vision` block work: a generated 720x360 image of large digits is read correctly, and a 3-red-circles/2-blue-squares image is counted correctly. (Note the default `VISION["cpu"]["max_tokens"] = 300` is below the 1024 the encoder itself warns about — fine for coarse tasks, worth knowing for grounding.)

So on Windows+AMD the encoder looks buildable from source today; only the packaging and the two guards above stand in the way.

## 6. One deliberate deviation worth a look

For `--vision cpu` I did **not** pass `--vram-reserve-mib`, although setup appends `VISION[vision]["reserve_mib"]` (700) for any non-none vision. In `generate.cpp`, `vram_reserve_mib` already defaults to 700, and passing it explicitly sets `vram_reserve_given = true`, which disables the automatic lowering of the reserve when the expert cache would get too small (#496). Since a CPU encoder uses no VRAM, keeping the automatic fallback is strictly better on a 12 GB card in low-RAM mode. If that reasoning is right, appending `--vram-reserve-mib` for `vision == "cpu"` may be unnecessary.

## Summary

* gfx1031 (RX 6800M, Navi 22) runs fine on Windows with a self-built HIP engine — same class of result as #524 (6700 XT on Linux) and #325 (9070 XT on Windows).
* Two real Windows bugs: `package_windows.py`'s missing encoding (build OK → package always fails on non-UTF-8 locales), and `build_vision_cpu()`'s `vcvars=None`.
* The Windows-HIP prebuilt path never installs a vision encoder, which is why Windows+AMD images are off — the encoder itself builds and runs.

Happy to send logs (`strata-coder-iq1_m.log`, build output) or test a patch if useful.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。