Issues / #1630

#1630 Windows HIP (RX 7900 XTX): a long prompt with `--prefill 8192` / `auto` makes the commit charge jump ~70 GB in one second, the engine dies at the commit limit (0xC0000409); 7500 and below stay flat

open · @rekaXua · 0 comentários · No GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Descrição

### Summary from Opus 5.5

On Windows with the ready-made HIP engine (0.1.41) on an RX 7900 XTX, the first long prompt read with 8192-token
prompt chunks (`--prefill auto`, which picks 8192 on this card, or `--prefill 8192` set by hand) makes the **system
commit charge jump by ~70 GB within one second**. `strata.exe`'s own private bytes do not change, so the allocation
is not visible in the process (driver / ROCm side?). The commit limit is then reached, Windows logs "Virtual Memory
Minimum Too Low", and the engine exits with `0xC0000409` (once `0xC0000005`); the server answers 503.

With smaller chunks the same prompts read fine and the commit charge stays flat. The threshold is between 7500 and
8192 on this PC.

Short prompts (which do not go through the chunked prompt path) are not affected, which is why this only showed up
with an agent client (OpenCode) sending ~20-30K-token first turns.

### Environment

- Windows 10 Enterprise LTSC 2021 (10.0.19044)
- GPU used: **AMD Radeon RX 7900 XTX 24 GB (gfx1100)**, HIP device 1, display driver 32.0.32015.2008
  (AMD Software: Adrenalin Edition 26.9.2, 26.20.15.02-260921a1)
- Second card in the PC (not used by this engine run): RX 6900 XT 16 GB (gfx1030), HIP device 0, driver
  32.0.21045.11001
- CPU: AMD Ryzen 9 5950X (16 cores, AVX2), RAM: 80 GB (79.9 GiB visible)
- Page file: SATA SSD 1: fixed 100,000 MB + SATA SSD 2: 4,096-35,840 MB; commit limit ~188 GB
- Strata v0.1.41 (`fb58e0d`), ready-made `strata-windows-x64-hip.zip` 0.1.41 (ROCm 10.2.0a20260930,
  hipBLASLt 100500), Python 3.11.9
- Model: OrcaRouter Qwen3.8-Flash-Next-Uncensored **IQ3_XXS**, packed per `docs/ORCA.md` (`iq_pack.py --compat-bf16`).

Engine args (only `--prefill` changed between runs):

```
--pack ...\packs\orca-iq3_xxs --native ...IQ3_XXS-00001-of-00002.gguf --ple-gguf ...IQ3_XXS-00001-of-00002.gguf
--expert-profile data\expert-profile.bin --expert-cache auto --prefill <N> --spec 4 --spec-min-p 0.5
--mtp ...\mtp\rt --max-context 262144 --kv int8 --kv-resident 32768
```

Config: `"gpu": 1`, `"backend": "hip"`, `STRATA_HIPBLASLT_TUNING=gfx1100-hipblaslt-100500.txt`, `"fit_max_tokens": true`.

### Measurements

Same request each time: one chat completion with a 26,902-token prompt (`docs/DETAILS.md` text as the system message),
`max_tokens` 64, `reasoning_effort` none, fresh engine per row. Commit charge and `strata.exe` private bytes sampled
every second (`Win32_OperatingSystem`, `Get-Process`); the watcher killed `strata.exe` above 140 GB commit so the
PC stayed usable.

| `--prefill` | prompt path borrow | commit before -> peak | `strata.exe` private | result |
|---|---|---:|---:|---|
| `auto` (log: `prompt chunk auto: 8192 tokens, a 96-slot ring`) | 1782 slots (3.61 GiB) | 98.8 -> **168.3 GB** in 1 s | 69.2 GB, unchanged | killed / crashes without the watcher |
| `auto`, no `STRATA_HIPBLASLT_TUNING` | 1782 slots | 98.8 -> **168.2 GB** | unchanged | killed |
| `8192` | 1782 slots (3.61 GiB) | 98.7 -> **168.3 GB** | unchanged | killed |
| `7500` | 1657 slots (3.36 GiB) | 98.7 -> 99.2 GB | unchanged | OK, 748 tok/s prompt read |
| `6144` | 1411 slots (2.86 GiB) | 98.8 -> 99.2 GB | unchanged | OK, 645 tok/s |
| `4096` | 1041 slots (2.11 GiB) | 98.9 -> 99.4 GB | unchanged | OK, 535 tok/s |
| `2048` | - | 98.8 -> 99.0 GB | unchanged | OK, 329 tok/s |
| `512` | - | 98.4 -> 98.7 GB | unchanged | OK, 173 tok/s |

A 66,235-token prompt also passed at 6144 (708 tok/s) and 7500 (782 tok/s) with a flat commit charge.

**The RX 6900 XT (gfx1030) in the same PC does not show it.** Same engine and model, the card's own config
(`"gpu": 0`, `--max-context 65536`, no hipBLASLt table), `--prefill 8192`: the prompt path borrows 1687 slots
(3.42 GiB) and the commit charge stays flat (90.8 -> 91.7 GB at 26,902 tokens, 91.5 -> 91.9 GB at 54,578 tokens),
356-370 tok/s prompt read (114 tok/s at `--prefill 512`). So it looks specific to gfx1100 (or to its driver
branch: the two cards run different driver builds, see Environment).

The jump happens within one sampling interval right after the prompt starts, i.e. at the first 8192-token chunk.
Disabling the hipBLASLt tuning table does not change it.

### What the logs show

Engine log (`strata-orca-7900xtx.log`) right after `2357 MiB of VRAM free with everything loaded`, with no request
line in between:

```
rocblaslt error: Could not initialize Tensile host: No devices found

rocblaslt error: Cannot read "D:\\Strata\\engine\\rocm\\bin\\hipblaslt\\library\\TensileLibrary_lazy_.dat" (or .zlib variant): No such file or directory

rocblaslt error: Could not initialize Tensile host:
Error 719(hipErrorLaunchFailure) C:/home/runner/_work/rockrel/rockrel/rocm-libraries/projects/hipblaslt/library/src/amd_detail/rocblaslt/src/tensile_host.cpp:3045:
hipGetDeviceCount(&count)
unspecified launch failure

rocBLAS error: Could not initialize Tensile host: No devices found
```

(These lines look like a consequence of the device being lost, not the cause: they appear after the commit jump.)

Windows event logs:

- System, Application Popup (ID 26): "Windows - Virtual Memory Minimum Too Low : Your system is low on virtual
  memory. Windows is increasing the size of your virtual memory paging file." - 15 s before the first crash.
- Application, ID 1000: `strata.exe`, faulting module `ucrtbase.dll` 10.0.19041.789, exception `0xc0000409`
  (twice), and once `Dbghelp.dll`, `0xc0000005`.

Other lines from every start, for context:

```
strata generate: expert arena: locked 14877 MiB via working-set minimum + VirtualLock; cudaHostRegister of the whole arena FAILED (invalid argument); 34 slices pinned (35 GiB); large pages refused for 53481570304 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata: Windows budgets 23748 of this card's 24560 MiB for this process; free VRAM is counted within that (STRATA_WDDM_BUDGET=0: off)
strata generate: expert cache 7289 slots, 14.77 GiB of VRAM
strata generate: KV streaming: 32768 of 262144 cells per QSA layer in VRAM, the K/V in 3.09 GiB of pinned RAM
strata serve: 2357 MiB of VRAM free with everything loaded
```

### Workaround

`--prefill 7500` (or 6144 for margin) in the config's args. With `auto` the engine picks 8192 on this card, so a
default setup on a 24 GB Windows AMD card may hit this on the first long prompt.

### Not checked yet

- Whether the ~70 GB is a fixed size or scales with the chunk / context (512K and 1M contexts were not run with
  long prompts at 8192).

Related: #1607 (engine dies silently at the Windows commit limit), #1130 (commit limits after allocation failure),
#1520 (Windows prefill staging and RAM).

No site

Links install, modelos, releases.