Pull requests / #876
#876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)
closed · @Hardin22 · 0 comentarios · En GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Descripción
In the resident RAM mode (`--resident-experts`), an adaptive round runs on a thread that the decode loop joins after the draft: 1. choose the swaps; 2. copy the evicted slots back with a stream sync; 3. copy the new experts in from the RAM copy. With ~90 swaps that cost us about 24 ms every 4th window. `--adapt-async 1` (opt-in, default off) splits a round into four steps that advance between windows on a helper thread, so no window waits for one: 1. **Choose and copy back.** Pick the swaps (`adapt()`'s rule, on snapshots of the counts and the residency) and copy the evicted slots back (D2H), each from the cache of the card that owns its layer. 2. **Stage and copy in.** Stage the exchanges, mark the evicted experts non-resident (the CPU reads them from the exchange buffers), and copy the new experts in, straight from a page-locked RAM copy or through a pinned bounce buffer. 3. **Mark resident.** Mark the new experts resident, and move the evicted blobs into the RAM places the new ones left. 4. **Flip** the copy's offsets. **Invariants:** - A slot is overwritten only after its old expert stopped being planned for the GPU, and a RAM place is reused only once its expert is resident. - The device residency tables are uploaded at steps 2 and 3, as `apply_pending` does, because the zero-doorbell graph reads them. - A request starts with the round drained, before its prompt lends any slot. It's off by default because it is not bit-exact from run to run: which window first computes a swapped-in expert on the GPU depends on when its copy lands. It falls back to the blocking tier, saying so once at startup, outside `--serve`, without the resident RAM copy, and with `--batch` slots, `--peer-device` or `--remote-expert-opt`. On a layer split, each swap's copies run on the owning card's refill stream; the homes are looked up as in #848's `resident_stage_swaps`. ## Measured Swift 1.5 IQ3_XXS at 160K (q4_0 KV), i9-14900KF, 32 GB of RAM, Windows 11, greedy decode, tok/s. Blocking tier → `--adapt-async 1`: | setup | blocking | async | |---|---|---| | RTX 4060 Ti + RTX 5080 split, resident mode (#848), `--trim-stage-weights`, stock draft layer, 5 prompts × 2 rounds, two interleaved pairs | 79.7 / 78.8 | 87.6 / 86.2 (**+10%**) | | same split, a fine-tuned draft layer and `--spec-min-p 0.65`, 4 prompts × 3 rounds | 81.0 / 81.8 | 108.1 / 107.0 (**+32%**) | | RTX 5080 alone, `--resident-experts` | 26.7 / 26.6 | 29.5 / 28.5 (noisy) | ## Checks - **Without the flag**: the greedy output equals main's byte for byte (4/4 prompts, 500 tokens, `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`). - **With the flag on the split**: 22 varied requests at temperature 0.6 (thinking on/off, five languages, code) swapped in 43,825 experts over 622 rounds. No degenerate answer, no refill error. ## Notes - **With #859** (`--pipeline-windows`): the two need fences between the async steps and the windows in flight. If both land as they are, `--pipeline-windows` has to go on this PR's off-list (`adapt_async_off`, one line); otherwise the pipelined loop's own fenced tier would run beside the async one. I have the proper combination working on a branch. I'll open it after this and #859, since it touches both. - **Cost**: the exchange buffers double (the bounce half): 96 more blob-sized pinned buffers, ~130 MiB with 1.38 MB blobs. - **Not tested**: HIP/SYCL.
En el sitio
Enlaces a install, modelos, releases.