贡献 / #1752

#1752 HIP: unbuffered file-tier reads and a locked expert arena (#1691)

open · @KevinX8 · 0 评论 · 去 GitHub 看

BenchmarksMulti-GPUAMD / HIP

说明

Bisected the gfx906 prompt stalls and 0.1.41 verify timeouts to 805088d (full story in #1691).

**What happens without this, on 1x MI50 16 GB (gfx906), ROCm 10.0 and 10.1, 60 GB RAM, 43 GiB resident budget:**

- 805088d's 5%-of-read-bytes rule puts this box on cached file-tier reads; 0.1.40 required room for all read bytes, so it read unbuffered.
- Cached here means ~42 GB of storage I/O per prompt through a ~6.6 GB cache (3.7x amplification in the engine's own file-tier counters) - relentless reclaim.
- The 43 GiB arena is HMM-tracked but unpinned (VmPin 0 after hipHostRegister), so reclaim unmaps it and KFD suspends ALL GPU queues: up to 45 of 60 s idle, 7-26 s per prompt. Past 60 s the watchdog kills the engine; starved verify spin flags time out instead.
- Plain v0.1.40.4: first prompt dies, ~34 tok/s limping. Plain v0.1.40 (same box/config): 358-366 tok/s, 0.0 s suspended.

**This PR (HIP-only, 11 lines, one file):**

1. Keep the old room >= read_bytes rule on HIP (the #1194 5% rule stays everywhere else - your P100 measurement stands where reclaim has no queue side effects).
2. Lock the whole HIP-registered arena (mlock), as hardening: it stops the crashes alone but not the 2x slowdown, the cache churn itself costs that.

**Measured on the same box, v0.1.41 + this PR, full benchmark (idle, 3x ~4k prompts, 300-token story, 6 GiB RAM pressure, prompt after):**

| | before | after |
|---|---|---|
| prompt read | ~34 tok/s (then dies) | 316-366 tok/s |
| story generation | n/a | 27.3 tok/s |
| GPU suspended per prompt | 7-26 s | 0.0-0.1 s |
| watchdog stalls / verify timeouts | 1+ per session | 0 |

Reverting only 805088d on plain v0.1.40.4 gives the same numbers, so (1) is the fix and (2) is belt and braces. Happy to split them into two PRs if preferred.

本站相关内容

相关页面的快捷入口。