Pull requests / #876

#876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)

closed · @Hardin22 · 0 comentários · No GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows

Descrição

In the resident RAM mode (`--resident-experts`), an adaptive round runs on a thread that the decode loop joins after the draft:
1. choose the swaps;
2. copy the evicted slots back with a stream sync;
3. copy the new experts in from the RAM copy.

With ~90 swaps that cost us about 24 ms every 4th window.

`--adapt-async 1` (opt-in, default off) splits a round into four steps that advance between windows on a helper thread, so no window waits for one:
1. **Choose and copy back.** Pick the swaps (`adapt()`'s rule, on snapshots of the counts and the residency) and copy the evicted slots back (D2H), each from the cache of the card that owns its layer.
2. **Stage and copy in.** Stage the exchanges, mark the evicted experts non-resident (the CPU reads them from the exchange buffers), and copy the new experts in, straight from a page-locked RAM copy or through a pinned bounce buffer.
3. **Mark resident.** Mark the new experts resident, and move the evicted blobs into the RAM places the new ones left.
4. **Flip** the copy's offsets.

**Invariants:**
- A slot is overwritten only after its old expert stopped being planned for the GPU, and a RAM place is reused only once its expert is resident.
- The device residency tables are uploaded at steps 2 and 3, as `apply_pending` does, because the zero-doorbell graph reads them.
- A request starts with the round drained, before its prompt lends any slot.

It's off by default because it is not bit-exact from run to run: which window first computes a swapped-in expert on the GPU depends on when its copy lands. It falls back to the blocking tier, saying so once at startup, outside `--serve`, without the resident RAM copy, and with `--batch` slots, `--peer-device` or `--remote-expert-opt`. On a layer split, each swap's copies run on the owning card's refill stream; the homes are looked up as in #848's `resident_stage_swaps`.

## Measured

Swift 1.5 IQ3_XXS at 160K (q4_0 KV), i9-14900KF, 32 GB of RAM, Windows 11, greedy decode, tok/s. Blocking tier → `--adapt-async 1`:

| setup | blocking | async |
|---|---|---|
| RTX 4060 Ti + RTX 5080 split, resident mode (#848), `--trim-stage-weights`, stock draft layer, 5 prompts × 2 rounds, two interleaved pairs | 79.7 / 78.8 | 87.6 / 86.2 (**+10%**) |
| same split, a fine-tuned draft layer and `--spec-min-p 0.65`, 4 prompts × 3 rounds | 81.0 / 81.8 | 108.1 / 107.0 (**+32%**) |
| RTX 5080 alone, `--resident-experts` | 26.7 / 26.6 | 29.5 / 28.5 (noisy) |

## Checks

- **Without the flag**: the greedy output equals main's byte for byte (4/4 prompts, 500 tokens, `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`).
- **With the flag on the split**: 22 varied requests at temperature 0.6 (thinking on/off, five languages, code) swapped in 43,825 experts over 622 rounds. No degenerate answer, no refill error.

## Notes

- **With #859** (`--pipeline-windows`): the two need fences between the async steps and the windows in flight. If both land as they are, `--pipeline-windows` has to go on this PR's off-list (`adapt_async_off`, one line); otherwise the pipelined loop's own fenced tier would run beside the async one. I have the proper combination working on a branch. I'll open it after this and #859, since it touches both.
- **Cost**: the exchange buffers double (the bounce half): 96 more blob-sized pinned buffers, ~130 MiB with 1.38 MB blobs.
- **Not tested**: HIP/SYCL.

No site

Links install, modelos, releases.