Pull requests / #920
#920 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
closed · @YamineRL · 0 comments · View on GitHub
Setup & installServer & APIAMD / HIPModels & quantsDocumentation
Description
Every IQ3_XXS run on my RX 6800 died the same way. Either `verify: timed out at layer N` (#267), or the serve watchdog's "no progress for 60 s". It happened on every run.
The box: RX 6800 16 GB (gfx1030), ROCm 7.2.4, 31 GiB RAM, Qwen3.8 Flash Next IQ3_XXS with `--mmap-experts`. Strata runs as a systemd user unit with `MemoryMax=28G`.
The fix is one line in the config's `env`:
```json
"env": { "GPU_PINNED_MIN_XFER_SIZE": "1048576" }
```
Since 2026-10-04 17:34 I've had zero stalls on IQ3_XXS.
Why it works is a guess, so treat it as one. For large copies from pageable memory, HIP may pin the source pages through a KFD userptr. With `--mmap-experts`, those pages are the mmapped expert file. When the kernel reclaims those pages, the GPU queues may be paused while it does. That would look exactly like a verify window that never finishes. Raising the threshold appears to stop HIP from pinning those copies. I haven't confirmed the mechanism in the HIP source or with a trace.
One related gotcha: don't put `MemoryHigh` on the unit. It counts page cache, which `--mmap-experts` lives on. A 16K prompt sat in `pread` for 8 minutes until the watchdog killed it. Use `MemoryMax` only.
This PR is docs only. If you want, I can follow up by having setup write the variable into `env` on gfx1030 when `--mmap-experts` is on (it would join `SETUP_ENV` in `setup.py`). I've left that out until you say you want it. I've only tested one card and one model.
Possibly related: #649 (gfx1030 verify timeouts; this is something to try alongside `STRATA_VERIFY_TRACE=1`), and #541 and #613 (gfx1201 prompt stalls, a different card, so not tested).
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.