Issues / #1706

#1706 Windows HIP: a fixed --expert-cache fails hard when other programs hold VRAM (0.1.40.2 3x slower, 0.1.41 launch failure)

open · @jerem91150 · 0 comentarios · En GitHub

Setup & installAMD / HIPModels & quantsWindows

Descripción

Native Windows 11, RX 7900 XTX 24 GB, 64 GB RAM, Qwen3.8-Flash-Next IQ3_XXS, `--expert-cache 7859` (fixed, calibrated on an idle card), 131K context.

Context: this came out of work I do across Strata and Project Maya on the same machine (Strata serves Qwen here, Maya runs GLM-5.3, and I compare and tune both). The VRAM-holding trick below served both: I wrote it for this Strata test and reused it on Maya.

With ordinary desktop programs open (Steam, a browser, Discord, a game launcher), about 4 to 5 GB of the card are taken before Strata starts. A fixed cache does not see that:

| same test set (2 scene prompts, 9K, 27K) | idle card | ~4.5 GB held by another program |
|---|---|---|
| 0.1.40.2, fixed cache | 81 s | 224 s with the real programs open (27K prompt 28.5 -> 126 s), Windows spills to shared memory |
| 0.1.41, fixed cache | 77 s | does not answer: `0 MiB of VRAM free`, then `iq_embed_rows: unspecified launch failure`, HTTP 503 |
| 0.1.41, `--expert-cache auto` | **74.6 s** | **80.4 s** |

To reproduce the "other program" without opening anything, I held 4.5 GB with a few `hipMalloc` + `hipMemset` calls from Python (ctypes on the bundled amdhip64_7.dll).

`auto` is also faster on the idle card here (27K prompt 28.3 -> 23.7 s): the room it leaves goes to the prompt path. So on Windows I would suggest that setup write `auto`, or that a fixed value be clamped at start when the free VRAM it measures is short, instead of the warning followed by the launch failure. The warning in 0.1.41 does name the cause and `auto`, which is how I found it.

En el sitio

Enlaces a install, modelos, releases.