Pull requests / #158

#158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUs

closed · @MrRoza · 0 comments · View on GitHub

Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux

Description

> **Note:** this PR was written by Claude (Anthropic's AI coding assistant, via Claude Code), working with @MrRoza on their own machine. The diagnosis, patches and measurements below come from that session; please review with that in mind.

## Problem

Under WSL2 (e.g. Docker Desktop on Windows), the NVIDIA driver can pin or map only about **1 GiB of host memory in total**. When it runs out, the WSL kernel logs:

    misc dxg: dxgk: create_existing_sysmem: establish_gpadl failed: -122

With the secondary expert caches (`--expert-cache-device1 5000 --expert-cache-device2 5000`) on 3 GPUs, startup failed twice because of this:

1. **CUDA2's context could not be created:**
   `strata generate: CUDA2 experts: cudaSetDevice(2) failed: out of memory (CUDA error 2)`
   CUDA1's context is created early (before CUDA0's weights, per the "proven startup order" comment), but CUDA2/3 are created after CUDA0's weights and MTP load. By then CUDA0's mapped host buffers (such as the 260 MiB native embedding) have used up the budget, and creating a new context needs some of it too.
2. **With that fixed, the native embedding couldn't be pinned:**
   `strata generate: native embedding: cannot pin 260 MiB`
   Three contexts now fit, but the mapped embedding table no longer does.

## What we checked

- Hardware: RTX 5070 Ti + 2x RTX 5060 Ti (16 GB each), Ryzen 9 5900X, 128 GB RAM, driver 616.92, Docker Desktop / WSL2, Swift 1.5 IQ3_XXS, 262144 ctx.
- The same config runs fine on **native Windows** with the prebuilt engine, so it's WSL-specific.
- In WSL2, 5070 Ti + **one** helper works; adding the second helper fails at `cudaSetDevice(2)` every time.
- `dmesg` inside the WSL VM shows hundreds of `establish_gpadl failed: -122` lines at startup. That's a limit on pinned/mapped memory, not free RAM: the VM had >80 GB free, and dropping caches doesn't help.
- The layer-split mode fails under WSL2 for the same reason (a third context OOMs; with 2 GPUs, the prefill hand-off `cudaHostAlloc` fails even at `--prefill 2048`). This PR doesn't change the layer split.

## The fix

1. **`STRATA_EARLY_REMOTE_CONTEXTS=1`** (opt-in; default behaviour is unchanged): creates *every* secondary context in the early block where CUDA1's already is, before CUDA0 allocates its weights, MTP and mapped buffers. The later loop skips contexts that already exist. Context creation then comes out of the budget first, while it's still available.
2. **Embedding VRAM fallback:** if `cudaHostAlloc` for the native embedding fails, the table is copied into the current device's VRAM (`cudaMalloc` + `cudaMemcpy`) and a line is logged, instead of failing the start. It's only ever gathered from, so a device copy works the same way; it costs its size in VRAM (260 MiB) and reads faster than over PCIe. The destructor frees whichever copy was used. Behaviour doesn't change when pinning succeeds.

## Result

With both changes, the 3-GPU helper setup starts and serves in Docker/WSL2:

    strata generate: CUDA1 context ready, 14.77 GiB free before expert arena registration
    strata generate: CUDA2 context ready, 14.77 GiB free before expert arena registration
    strata: native embedding: cannot pin 260 MiB, kept in VRAM instead
    strata generate: CUDA1: 5000 additional experts, 8.10 GiB; results return through pinned host rows
    strata generate: CUDA2: 5000 additional experts, 8.10 GiB; results return through pinned host rows

Under WSL2, the second helper doesn't make it faster than one helper (39K-token cold prompt: 625 vs 636 t/s prefill, 41–50 vs 46–53 t/s decode). Presumably the pinned-memory limit also slows the helpers' result path. On native Windows the same 3-GPU setup does 944 t/s prefill and 47–60 t/s decode. So this PR is about **starting** reliably under WSL2, not speed.

Tested on engine 0.1.24 (commit 3ce2523). The branch is rebased onto current `main` (8fc40dd), where both changes apply cleanly, but that rebased build hasn't been run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.