Issues / #771

#771 Linux expert arena (0.1.39): MADV_HUGEPAGE + defrag=madvise makes the arena's faults 20x slower (25 s start -> 432 s), and 'loaded ... at X GiB/s' does not show it

closed · @Alucard24 · 6 Kommentare · Auf GitHub

Server & APINVIDIA / CUDAModels & quantsLinux

Beschreibung

> **Update (measured after this report).** The 3.06-3.12 GiB/s load line above is not what this bug costs.
> The same profile that crawled here reached `ready` at **432 s** while still printing `loaded 39.97 GiB at 2.86 GiB/s`,
> and the same profile with the MADV_HUGEPAGE request skipped is ready in **~25 s**. The isolated 20x `madvise`
> measurement, the /proc numbers, the switch and the note change are in the comment below.

### Symptom

Engine 0.1.39 (source build), arena profile (`strata-iq3_xxs.json`: IQ3_XXS native pack, experts in RAM, no
`--mmap-experts`, `--max-context 131072 --kv-resident 32768`, MTP and vision on).
Machine: Manjaro Linux, kernel 6.12.108-1-MANJARO, i7-11700B, 62 GiB RAM, RTX 5070 Ti (16.3 GiB).

With the RAM free the fill is fast, and the log names the request that is in play:

```
strata generate: expert arena: cudaHostRegister PORTABLE ok; MAP_HUGETLB unavailable (needed 20464 2 MiB pages, vm.nr_hugepages=0); transparent huge pages requested (MADV_HUGEPAGE)
strata generate: loaded 39.97 GiB at 3.12 GiB/s
```

(a later start of the same profile, same engine: `strata generate: loaded 39.97 GiB at 3.06 GiB/s`)

With the RAM already in use, the same 39.97 GiB fill crawls at **60-77 MB/s**, one core at 100 %, no disk I/O,
and the PC becomes unusable while it fills. Two starts of the arena profile did that in one afternoon; both were
stopped, and the engine's output is buffered, so those logs end at the line *before* the arena note:

```
strata generate: GPU 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0     <- last line in those logs
```

The numbers below are live `/proc` readings taken while the fill was in progress, not log lines:

- RSS deltas: +484 MB in 8 s, +616 MB in 8 s (60-77 MB/s); RSS 31.7 GB of ~43 GB when sampled
- `/proc/vmstat`, 8 s window: `compact_stall +858`, `compact_fail +858`, `thp_fault_alloc +0`,
  `thp_fault_fallback +290`, `allocstall_normal +286`, `nr_free_pages -498 MB`
- the process: 18,304 minor faults/s; at 4 KB per page that is 75 MB/s, the same number the RSS deltas gave.
  CPU 95-100 % of one core. In the same window the device read 0-1 MB/s and wrote nothing.
- what the user sees: the server repeats `[strata] still starting (N s) - please wait ...` (serve/server.py:307)

### Why the `MADV_HUGEPAGE` request is my prime suspect

This machine: `/sys/kernel/mm/transparent_hugepage/enabled = [always]`, `defrag = [madvise]`,
`vm.nr_hugepages = 0`.

The kernel's own description of the defrag policies (Documentation/admin-guide/mm/transhuge.rst):

- `madvise`: "will enter direct reclaim like always but only for regions that are have used madvise(MADV_HUGEPAGE). This is the default behaviour."
- `always`: "... an application requesting THP will stall on allocation failure and directly reclaim pages and compact memory in an effort to allocate a THP immediately."

With `defrag=madvise`, the arena's own `madvise(MADV_HUGEPAGE)` is what can put this 40 GiB mapping into
synchronous direct reclaim and compaction in the faulting thread. The counters point the same way: in that 8 s
window a compaction was tried ~107 times per second and never succeeded (`compact_fail` == `compact_stall`), no
new THP was allocated (`thp_fault_alloc +0`), the faults fell back to 4 KB pages (`thp_fault_fallback +290`), and
at 18,304 faults/s the faulting thread paid ~50 µs of kernel work per page - 75 MB/s with one core burnt.
(The arena does get huge pages earlier in a run: 13.8 GB of a 31 GB fill were backed by them when sampled.)

With the RAM still free there is nothing to compact, and the same request costs nothing (3.06-3.12 GiB/s above).
For reference, the 0.1.38 build (which has no THP request: its note ends in `; using 4 KB pages`) filled the same
profile on this machine at 3.32 GiB/s, so the request buys nothing in this phase here; the TLB argument in the
comment above it is about the CPU expert pool during decode, not this fill.

### Not the cause (checked)

- disk: the device did not move during the window (0-1 MB/s reads, no writes); the expert bytes come from the
  page cache and the pack
- loader/pack: the same code path, same pack and same machine filled the same arena at 3.06-3.12 GiB/s minutes
  earlier
- VRAM: not used in this phase (the fill is host memory; the cache slots are chosen afterwards)
- an OOM: no swap activity and no OOM kill, the fill kept progressing - the pages it got were simply 4 KB ones
- `STRATA_NO_LARGEPAGES=1` is not a workaround: it also skips the hugetlb attempt, i.e. it changes two things

### Same mechanism reported elsewhere

- golang/go#61718 "runtime: MADV_HUGEPAGE causes stalls when allocating memory"; the Go runtime dropped its own
  huge-page policy in 1.21.4 ("the Go runtime would no longer try to impose a huge page policy itself")
- cockroachdb/cockroach#130241 "server: provide guidance and/or software control over transparent huge pages"
- coreos/bugs#2635 transparent huge pages set to [always] are sub-optimal for many applications

### What I would propose (implemented locally, with a selftest - happy to open a PR)

1. put the policy in the note, so a crawling start explains itself in the log:
   `...; transparent huge pages requested (MADV_HUGEPAGE; kernel defrag=madvise)`
   (one read of `/sys/kernel/mm/transparent_hugepage/defrag`; the note is what users paste into reports)
2. a same-run A/B switch for the request only: `STRATA_NO_ARENA_THP=1` ->
   `; transparent huge pages skipped (STRATA_NO_ARENA_THP); using 4 KB pages`. Default unchanged.
   Alternative if you prefer no new switch: skip the request, or warn, when
   `/sys/kernel/mm/transparent_hugepage/defrag` is `always`, `madvise` or `defer+madvise` - i.e. whenever a
   2 MiB block may be impossible to get synchronously.
3. `src/core/pinned_thp_test.cpp`: a selftest for the note in both cases (64 MiB arena, no model, no 40 GiB).

If you want a different default than 2, say so and I will follow that instead. I can also report a before/after
arena run with the switch on this machine, but that is the run that makes the PC unusable here, so it is a
one-shot: I would rather do it once you tell me what direction you want.

Mehr auf der Site

Links zu Install, Modellen, Releases.