Issues / #1250

#1250 [Linux][RTX 3080 12 GB, 62 GB RAM] 0.1.40.1 start OOMs the host while allocating the page-locked expert copy (IQ2_XS); 0.1.31 starts fine

open · @Absolutely-not-a-coder · 6 commentaires · Sur GitHub

Server & APINVIDIA / CUDAModels & quantsLinux

Description

## Summary

On a 62 GB Linux box with a 12 GB RTX 3080, **0.1.40.1 is OOM-killed during start** with IQ2_XS at 262K context. The same model, config and machine start and serve fine on **0.1.31** (service peak 47.7 GiB). The low-RAM modes don't avoid it.

`--resident-experts`, `--resident-budget-gib 26` and `--resident-budget-gib 26 --pcie-frac 0` all reach `session is up (engine 0.1.40)` with the GPU cache filled. Each then dies while `FileExpertSource` allocates its page-locked cache complement. In those three runs it was a **host-wide** OOM (`global_oom`), even though `MemAvailable` was 54 GB a few seconds earlier.

## Environment

- Ubuntu 24.04.3, kernel 7.0.0-34-generic, NVIDIA driver 595.91.07
- GPU: RTX 3080 12 GB (sm_86), one card
- CPU: i5-14400F (AVX2, no AVX-512); RAM 62 GB; swap 8 GB (swappiness 60, overcommit 0)
- Build: `v0.1.40.1` (82f46a8), `STRATA_PORTABLE=ON`, `CMAKE_CUDA_ARCHITECTURES=86`, `GGML_NATIVE=OFF` + AVX2/FMA/F16C, CUDA 13.0.88, ggml 3cf0325. Runtime CUDA libs come from the venv's `nvidia/cu13` (13.0.96) via `lib_dirs`.
- Model: ISTA-DASLab Qwen3.8-Flash-Next-GSQ-RCO **IQ2_XS**, native pack made by `tools/iq_pack.py` (no `experts.bin`). Shard 1 (39.2 GB, experts) is on a SATA SSD (ext4); shard 2 (28.8 GB, PLE) is on NVMe. MTP draft enabled.
- Run as a systemd service: `MemoryHigh=54G MemoryMax=56G MemorySwapMax=0 LimitMEMLOCK=infinity OOMScoreAdjust=500`. Nothing else large runs on the host: about 2.5 GB anon outside the engine, with the GPU's other tenant stopped.

Engine arguments (from the server config):

```text
--pack <pack>/iq2_xs --native <shard1>.gguf --ple-gguf <shard2>.gguf --expert-profile data/expert-profile.bin
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <mtp>/rt
--max-context 262144 --kv int8 --kv-resident 32768 --vram-reserve-mib 1800
```

## What happens

| Run | Last engine line | Memory at the kill | Kill |
|---|---|---|---|
| defaults | after `GPU 0: NVIDIA GeForce RTX 3080, compute capability 8.6`, while "loading the experts into RAM (about 39 GB)" | anon ~25 GB, shmem 4 GB, page cache 27-31 GB | cgroup OOM at 56 GiB |
| defaults, `STRATA_READ_AHEAD=0` | same | same | cgroup OOM at 56 GiB |
| defaults, without `LimitMEMLOCK=infinity` | same | cgroup counted only 21.8 GB; process anon 29.5 GB + shmem 3.9 GB | **global** OOM |
| `--resident-experts` | `FileExpertSource: allocating 32.35 GiB page-locked cache complement` | Shmem 4 → 25.5 GB in ~5 s, MemFree 0.9 GB, Cached 53 GB | **global** OOM |
| `--resident-budget-gib 26` | `FileExpertSource: allocating 26.00 GiB page-locked cache complement` | Shmem → 19-22 GB, MemFree 0.9 GB, Cached 53.6 GB | **global** OOM |
| `--resident-budget-gib 26 --pcie-frac 0` | same | same | **global** OOM |

`/proc/meminfo` every 5 s (MiB), `--resident-experts` run:

```text
21:11:14 MemFree:16322 MemAvailable:54525 Cached:38249 Unevictable:4 Mlocked:0 AnonPages:2794 Shmem:4012
21:11:19 MemFree:11515 MemAvailable:50291 Cached:42830 Unevictable:4 Mlocked:0 AnonPages:2929 Shmem:8020
21:11:24 MemFree:1254  MemAvailable:32681 Cached:53230 Unevictable:4 Mlocked:0 AnonPages:2929 Shmem:25499
```

Kernel log, budget + `--pcie-frac 0` run:

```text
avahi-daemon invoked oom-killer: gfp_mask=0xcc0(GFP_KERNEL), order=0, oom_score_adj=0
oom-kill:constraint=CONSTRAINT_NONE,...,global_oom,task_memcg=/system.slice/<unit>.service,task=strata,...
Out of memory: Killed process (strata) total-vm:87225700kB, anon-rss:826992kB, file-rss:2366172kB, shmem-rss:26818796kB
```

The engine log before the kill (resident run) looks normal: `session is up (engine 0.1.40)`, `pre-filled 3250 of 3250 slots from the profile in 29.8 s`, `token graph hit path: 3250 resident experts`. Then:

```text
FileExpertSource: RAM budget 26.00 GiB: 19353 of the 28.65 GiB of experts the GPU cache does not hold, by profile rank; the rest are read from the files
FileExpertSource: allocating 26.00 GiB page-locked cache complement
```

## What it looks like (a guess, not verified)

The page-locked complement is committed as shmem much faster than the kernel gives back the ~30 GB of clean file cache left by reading the GGUF shards. Either that cache can't be reclaimed at that moment (mapped and in use?), or the pinned allocation path doesn't wait for reclaim. The 4 GiB headroom check (`STRATA_RESIDENT_HEADROOM_GIB`) passes because `MemAvailable` counts that cache as available, and it doesn't see a cgroup `memory.max` either.

Possible mitigations: allocate and pin the complement in chunks, letting reclaim catch up (or `posix_fadvise(DONTNEED)` the shard ranges already copied); size the budget against `memory.max` when running in a cgroup; and treat an allocation that can't be satisfied as "fall back to the mapped mode" rather than faulting pages until the OOM killer runs.

I can rerun with `STRATA_RSS_TRACE=1` or anything else that helps. This box also runs other services, so I'd rather not trigger more global OOMs blind.

Sur le site

Liens install, modèles, releases.