贡献 / #1601

#1601 Memory guard: parking and session save/restore check the RAM the container really has (`--memory-limit-mib`)

open · @noon-at-cgn · 0 评论 · 去 GitHub 看

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDADocumentationWindows

说明

## Title
Issue: Related: #1250 (host OOM while page-locking the expert copy; its `STRATA_PIN_GUARD` covers the start-up copy), #1572 / #1573 (SAVE admission, open draft).

## Summary
**Why it matters:** parking a conversation (about 240 MB to 750 MB per park) and saving or restoring a session file decide with `MemAvailable` alone. Inside a container that figure ignores the container's own limit, so the engine can start a copy it has no room for. In our log from a bad day, `MemAvailable` read 21.8, 25.4 and 3.7 GiB just before three kills while the container's cgroup was at 98.6, 94.9 and 96.6 GiB of a ~100 GiB limit. `available_host_bytes()` now returns the smallest of `MemAvailable`, the room under every cgroup limit above the engine, and an operator cap (`--memory-limit-mib N`, the total RAM the container may use). With no cgroup limit and no flag it is `MemAvailable`, as before. Those kills were host-wide OOMs (the host was over-committed), so the PR does not claim to prevent them; it makes the engine's own admission decisions use the real headroom.

| 90 GiB container, resident RAM mode | `--memory-limit-mib 101376` | `--memory-limit-mib 78000` |
|---|---|---|
| parking attempts that parked | every one (135 parks, 0 skips in one 308-request run) | 0 of 8 |
| log line | `parked N tokens in ... ms` | `skip parking (... 0 MiB available, source flag)` |
| requests served | normal | normal (decode 79.6 tok/s) |

The engine held 89.7-89.9 GiB, above a 76.2 GiB cap, so headroom is 0: the intended refusal, and the request goes on.

## What changed
- `conversation_memory.hpp/.cpp`: `conversation_available_memory()` becomes `available_host_bytes()` / `sample_host_memory()`. The cgroup term reads `memory.max` and `memory.high` (v2) or `memory.limit_in_bytes` (v1) of the process's group and each visible ancestor; clean inactive file cache is credited. An unparsable cgroup file makes the sample unknown, and unknown refuses.
- `src/program/generate.cpp`: flag and variable, one start-up line (limit, source, usage, credit); parking (before and after capture), save and restore read the new figure; a refusal says how much was available and from which source. The eviction loop is unchanged.
- `sycl/src/program/generate.cpp`: four callers renamed. `tests/core/memory_guard_test.cpp` (new, CPU, fixture directories for `/proc/meminfo` and the cgroup tree) + `CMakeLists.txt`; `docs/DETAILS.md`.

## Extra Notes
Measured on engine 0.1.40.3 (d5ea713) plus this change; the branch itself is cut from main fb58e0d.
Relation to #1250: not duplicated, not touched. `STRATA_PIN_GUARD` decides at start whether to page-lock in steps; this covers parking, SAVE and RESTORE after the start. Both read cgroup limits, in separate code (`conversation_memory.cpp` also builds standalone for the conversation tests); a shared reader can follow if you prefer one. #1573 edits the same SAVE lines, so one of the two needs a rebase.
Machine: 2x RTX 3080, 90 GiB container, UD-Q4_K_XL. Checked: `memory_guard_test` (130 checks; an existing binary, the rebuild was blocked by the build lock) and `conversation_memory_test`, CUDA build. Not built: HIP, SYCL, Windows. Not measured: a cap between 78000 and 101376, any run without the flag, any throughput effect.

本站相关内容

相关页面的快捷入口。