Pull requests / #80

#80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)

closed · @andrewcoul · 0 comentários · No GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Descrição

## Why

The docs say 32 GB of RAM is not enough: the arena keeps all 24,576 experts pinned in RAM. But once the profile has filled the VRAM cache, those experts' RAM copies are never read again (the cache does not evict, prefill skips resident experts, verify windows compute them from VRAM). On a 16 GB card that is ~5-7k experts, and the coldest ones in the profile are rarely routed at all.

`--tiered-experts` uses that. After the cache fill, every expert is sorted into a tier:

- **VRAM**: cached experts; their file pages are dropped from RAM.
- **Pinned**: the next experts by profile rank, up to `--host-budget-gib` (auto = `MemAvailable` minus `--host-reserve-gib`, default 8). The driver refuses `cudaHostRegister` on file-backed pages ("operation not supported"), so each run is remapped to anonymous memory in place (`MAP_FIXED`, same addresses, so `blob_offset` is unchanged), filled with `pread`, then registered.
- **Cold**: the rest stay in the mapped `experts.bin` and page in from the SSD. A routed cold expert is prefetched (`MADV_WILLNEED`) from helper threads so the CPU pool does not stall on it.

## Results (one machine)

Ryzen 7 7800X3D, RTX 4080 SUPER 16 GB, **32 GB RAM**, NVMe (~6.5 GB/s), Linux, ISTA-DASLab GSQ-RCO IQ3_XXS. Engine args: the usual ones plus `--tiered-experts --host-reserve-gib 14 --vram-reserve-mib 1500 --adapt-every 0 --spec-min-p 0.7`.

- Tiers at reserve 14 GiB: ~5.1k experts in VRAM, ~7.5k pinned (12 GiB), ~12k on the SSD (20 GiB).
- Decode, 400-token chat replies (MTP on): median ~53 tok/s (range 37-62) with an idle desktop; 44-46 median in a later session when other apps were using more RAM and VRAM, because fewer experts get cached. That dependence is by design.
- Prefill of a 2.3K-token prompt: 21 s in my first tiered version -> 6 s with the stager changes here. There is no upstream baseline to compare with: the default arena does not fit on this machine.
- Q2_0 at reserve 10 GiB: median ~56 tok/s.

## Other changes

- `ExpertCache::fill_slot` / `fill_slot_blocking`: a blob that straddles the end of a registered run makes `cudaMemcpyAsync` fail with "invalid argument"; falls back to a locked pinned bounce buffer. Serve-mode adaptive swaps now go through `fill_slot`.
- `ExpertSource::dma_capable()` replaces the `device_alias(layer, 0)` probe when deciding whether misses can be read over PCIe (with tiers, expert 0 of a layer may not be pinned).
- A native (IQ) pack writes `experts.bin` from shard 1 on first use of the tiered source.
- Prefill stager, **only** for a source that streams from the SSD (`streams_from_ssd()`): ring 16 -> 48, up to 8 threads (`STRATA_STAGER_THREADS`), and unpinned experts are read with `O_DIRECT` `pread`.
- Windows: `TieredExpertSource` compiles to a stub and `open()` reports "not available".

Defaults are unchanged without `--tiered-experts`.

## Testing and limits

- Built with CUDA 13.3 / sm_89 / g++-15 and run on that one machine, rebased on `main` (engine 0.1.20, c1e9033). It also builds on its own, without #79 (my numbers include #79).
- Old vs rebased build, same conditions: decode within run-to-run noise, 2.3K prefill 7.5 s vs 6.1 s. Serve mode with adaptive swaps on and off ran without refill errors.
- **Not run:** the default (non-tiered) path at runtime (the arena needs >40 GB here; it is compiled, and the changes there are the `dma_capable()` refactor and a `fill_slot` fallback that only triggers on `cudaErrorInvalidValue`); Windows (stub only, not compiled there); the parity tests (`-DSTRATA_BUILD_TESTS=ON` does not configure on `main`: `tests/` and several `bench/micro` sources are missing).

## Questions

- `--adapt-every 0` was clearly better in tiered mode in my A/B runs (34 -> 44 and 44 -> 49 tok/s), since each swap refills VRAM slots from cold experts. Should tiered mode turn it off by default?
- `setup.py` still refuses below 44 GB. I did not touch it; happy to add a tiered path (RAM check plus the flags in the generated config) if you want this in the installer.
- The reserve default (8 GiB) is a guess; on 32 GB, 10-14 measured best.

No site

Links install, modelos, releases.