Issues / #633
#633 The pinned expert arena is allocated without a host-RAM check: a container limit or a low MemAvailable ends in the OOM killer mid-load, not a message (the file tier already has the probe)
closed · @Avicennasis · 1 comments · View on GitHub
Description
On `main` at `99f3dbd` (0.1.38). Host RAM, not VRAM: #486 (the weight arena / a hard-frozen host at 262K) and #620 (the native head upload's out-of-memory at startup) are the VRAM side of the same shape; this is the pinned expert arena against the RAM the process can actually get.
## What is there
`available_memory_bytes()` (`src/core/expert_source.cpp:220-283`, in the file's anonymous namespace) reads `MemAvailable` from `/proc/meminfo`, then walks the cgroup-v2 ancestors of the `0::/` line in `/proc/self/cgroup` under `/sys/fs/cgroup`, taking the minimum of `MemAvailable` and each `memory.max` minus what `memory.current` / `memory.stat` say is held; on Windows it is `ullAvailPhys`. `FileExpertSource::pin_cache_complement` uses it for `--resident-budget-gib` (`:1263`): a budget over the room is clamped to `available - headroom - 256 MiB` and says so with its numbers (`:1270-1285`, #403), and a complement that still does not fit is refused with the numbers named (`:1326-1331`):
```
FileExpertSource: resident complement 28.50 GiB exceeds available RAM (30.10 GiB) minus the 4 GiB safety headroom
```
That is the right shape. It only runs on the file tier (`--mmap-experts`).
## What is not
`ArenaExpertSource::open` (`:2545`) - the pinned arena, the default tier - checks the pack's size against the layout before allocating (`:2560-2581`, "SIZE CHECK BEFORE THE ALLOCATION, not after") and then goes straight to `new PinnedArena(want + blob, ...)` (`:2597`). Nothing reads the host's available memory on that path: `generate.cpp:2783` calls `open` after the pin-cap decision, and the only `MemAvailable` read in `generate.cpp` (`:781`) feeds the stall report (`:811-813`), not a decision. On Linux the arena is an anonymous `MAP_PRIVATE` mmap (`src/core/pinned.cu:227`), so the mapping of 30-77 GB succeeds whatever the host has, the pages are committed as `load_experts_ranges` writes them, and a host (or a container with a `memory.max` under the arena) finds out from the OOM killer partway through the load, or from swap - the hard freeze #486 describes, from the other side. The message, when there is one, is the kernel's. A 128 GB host whose `MemAvailable` is 20 GB because another engine is still exiting gets the same ending.
Two smaller things on the probe itself:
- No cgroup v1. The walk only looks for the `0::/` line; on a v1-only host (older Docker hosts, RHEL 7/8 defaults, some Kubernetes nodes still) there is none, `resolved_v2` stays false and the function returns false, so the file tier refuses the budget with "cannot determine available RAM" rather than reading `MemAvailable` and `memory/memory.limit_in_bytes`. The arena, once it has the probe, should not inherit that.
- `src/platform/memory_test.cpp` (35 lines) only exercises `lock_resident` on 256 MiB; the cgroup walk has no test, so a fake root with `memory.max` = `max`, a number, a missing file, a malformed value, and a v1 tree cannot be checked without a container.
## What I would do
Move `available_memory_bytes` to `src/platform/memory.{hpp,cpp}` beside `lock_resident` (same file-local logic, a root path parameter so a test can point it at a fake tree, `/sys/fs/cgroup` by default), add the v1 fallback (the `memory` controller's path from `/proc/self/cgroup`, `memory.limit_in_bytes` and `memory.usage_in_bytes`; no cgroup line at all means `MemAvailable` alone, not "cannot determine"), and call it in `ArenaExpertSource::open` right after the size check, before the `PinnedArena` - the arena knows `want + blob` there. Refuse with the three numbers the way the file tier and the #486 VRAM message do, plus what makes room:
```
ArenaExpertSource: the expert arena needs 31.42 GiB but 20.10 GiB of RAM is available (11.32 GiB short) - --mmap-experts (with --resident-budget-gib for the hottest experts) runs with less RAM
```
`STRATA_RESIDENT_HEADROOM_GIB` (`generate.cpp:1313`) is the natural headroom to reuse. `memory_test.cpp` gets the fake-root cases: v2 `max` / numeric / missing / malformed (refuse), v1 limit, no cgroup line. The file tier keeps calling the same function, so its behaviour on v2 is unchanged and on v1 it stops refusing.
Size: M - the move and the call are ~40 lines, the v1 branch ~30, the tests the rest; nothing in the cache policy, the adaptive tier or the conversation cache. I can send it as a PR if you want it; filing first because it adds a refusal to the default start path and you may prefer a warning there, or a different headroom.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.