Issues / #1607

#1607 Windows: the engine dies without a message (0xC0000409) when the commit limit is reached; the RAM checks only look at physical RAM

open · @masel · 0 Kommentare · Auf GitHub

Setup & installNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Beschreibung

## Summary

On Windows, UD-IQ4_XS without `--resident-budget-gib` crashed 5 times in 12 minutes under an agentic workload, with
exit code 3221226505 (0xC0000409; once 3221225477, 0xC0000005) and no message from the engine. Each crash came when the
system's **commit charge** reached the commit limit (95.88-95.90 of 95.93 GiB), while 2.6-4.4 GiB of physical RAM were
still available. Windows logged "virtual memory insufficient" (System event 2004, Resource-Exhaustion-Detector) at the
same times. The same 107 requests replayed with WSL shut down (commit max 92.0 GiB) gave no crash.

The engine's RAM admission for the conversation cache and session checkpoints reads only `ullAvailPhys`
(`src/core/conversation_memory.cpp`, `conversation_available_memory()`), so on Windows it admits allocations that the
commit limit then refuses. The commit limit is RAM + page file, and under WDDM the card's VRAM is charged to it too
(#141), so a page file that setup accepts (it warns below 4 GB, #60) can leave almost no commit headroom.

## Setup

- RTX 5060 Ti 16 GB (driver 616.92), Ryzen 7 5700X3D, 80 GB RAM (79.9 GiB visible), Windows 11 Pro 10.0.26200
- Page file fixed at 16 GB -> commit limit 95.93 GiB
- Strata 0.1.40.3 (engine and server), release build
- Model: unsloth UD-IQ4_XS (55.4 GiB of experts), config as setup writes it but without `--resident-budget-gib`
  (see #1080): `--max-context 393216 --rope-scaling yarn --rope-scale 1.5 --kv int8 --kv-resident 32768
  --expert-cache auto --spec 4 --pcie-frac 0.20 --spec-min-p 0.70`; `STRATA_TRACE=1 STRATA_RSS_TRACE=1` in `env`
- Load: OpenCode 1.18.35 in a Docker container in WSL 2, three coding tasks (8-30K-token conversations, tool calls),
  each run twice; requests reach the server through a small TCP bridge on 127.0.0.1

## What was measured

Commit charge, commit limit, available RAM, strata.exe working set and private bytes and the WSL VM's working set,
every 10 s (`Win32_PerfFormattedData_PerfOS_Memory`, `Get-Process`).

| Phase | Requests | Engine crashes | Commit max | Available RAM min |
|---|---:|---:|---:|---:|
| A: agentic, WSL + Docker running | 107 | **5** | 95.92 of 95.93 GiB | 2.6 GiB |
| B: the same 107 request bodies replayed from Windows, WSL shut down | 107 | **0** | 92.0 of 95.93 GiB | 6.1 GiB |

The samples just before each crash in phase A (time, available RAM, commit, strata.exe private bytes, WSL VM working set):

```
15:37:23  2.93 GiB  95.84 GiB  79.02 GiB  3.32 GiB   -> crash, 0xC0000409
15:41:30  3.21 GiB  95.89 GiB  79.05 GiB  2.98 GiB   -> crash, 0xC0000409
15:43:24  4.38 GiB  95.88 GiB  78.79 GiB  1.91 GiB   -> crash, 0xC0000005
15:47:00  4.01 GiB  95.88 GiB  79.15 GiB  1.98 GiB   -> crash, 0xC0000409
15:49:24  3.96 GiB  95.90 GiB  79.04 GiB  1.98 GiB   -> crash, 0xC0000409
```

System event 2004 in phase A at 15:16:36, 15:22:59, 15:27:59, 15:33:10, 15:39:19, 15:43:16, 15:46:23, 15:49:04,
15:51:35 and 16:15:52 (Windows raises it once the commit stays near the limit; not every one ended in a crash, none in
phase B). Its text, for example:

```
Windows hat diagnostiziert, dass der virtuelle Speicher unzureichend ist. Die folgenden Programme belegten den meisten
virtuellen Speicher: strata.exe (21016) belegt 84826337280 Bytes, vmmemWSL (4584) belegt 4080967680 Bytes und
claude.exe (5908) belegt 623763456 Bytes.
```

(That one is from the first occurrence, on the previous day with engine 0.1.40: five crashes with the same exit code in
the same kind of run, event 2004 ten times between 18:14 and 19:10.)

strata.exe's private bytes start at ~69 GiB after loading and grow to ~79 GiB over the requests (conversation cache,
checkpoints), so the headroom shrinks during a session; a freshly restarted engine is back at 69 GiB, which is why the
crashes come in a cluster and then stop for a while.

## Where it dies

Every crash came right after a prompt read, before the first decode step, which is where the checkpoints are saved.
The last engine log lines before each restart (STRATA_TRACE):

```
strata trace: prompt chunk 0 of 1954
strata trace: read 1954 tokens (batched) in 7923.8 ms
```
```
strata trace: read 16 tokens (windows) in 377.0 ms
```
```
strata verify: capturing the 6-token window (295 MiB of VRAM free)
```

The engine log has no error line; the server reports `the engine stopped unexpectedly (exit code 3221226505)`.
0xC0000409 fits an allocation that throws where nothing catches it (`std::terminate` -> fail-fast). The conversation
cache's own `try/catch (std::bad_alloc)` in `generate.cpp` would have logged "allocation failed; skip parking", which
does not appear, so the failing allocation is elsewhere. Minidumps of all five crashes (~240 MB each, WER LocalDumps,
DumpType 1) are available if useful.

## Suggestions

- Engine, Windows: let `conversation_available_memory()` (and the session/checkpoint admission that uses it) return
  `min(ullAvailPhys, ullAvailPageFile)` minus the floor, so a snapshot is skipped when the commit cannot take it.
  `expert_source.cpp` already reads `ullAvailPageFile`.
- Engine: catch `std::bad_alloc` around the prompt-checkpoint saves as the conversation cache does, and install a
  `std::set_terminate` / unhandled-exception handler that writes one line ("out of memory: commit X of Y GiB") to the
  log before exiting, so this case is recognisable.
- Setup: the page-file warning (below 4 GB) does not catch this case. A check against the model's needs would: experts
  in RAM + VRAM + pinned KV + a margin for the conversation cache and the rest of the system, compared with RAM + page
  file. Here: 55.4 GiB experts + 16 GiB VRAM + 4.6 GiB KV = 76 GiB of 95.9 GiB before anything else runs.
- Docs: "System managed" or a page file of RAM/2 or more for the larger models on Windows.

## With a larger page file

Page file raised from 16 GB to 32 GB (commit limit 95.93 -> 111.93 GiB), reboot, phase A repeated unchanged (same
config, same three tasks twice, WSL + Docker):

| Page file | Engine crashes | Commit max | Available RAM min | strata.exe private max | Event 2004 |
|---|---:|---:|---:|---:|---:|
| 16 GB | 5 | 95.92 of 95.93 GiB | 2.6 GiB | 79.2 GiB | 10 |
| 32 GB | **0** | 93.96 of 111.93 GiB | 3.3 GiB | 79.1 GiB | 0 |

Same memory use by the engine, the same low physical RAM (3.3 GiB available), no crash: the limit that was hit is
the commit, not physical RAM. The page file's peak use since the reboot was 1.5 GB of 32 GB
(`Win32_PageFileUsage`); it mainly raises the limit.

Mehr auf der Site

Links zu Install, Modellen, Releases.