Pull requests / #1517
#1517 Adaptive tier: evictions reach the device's residency table before the next window when the tier does not wait (fix)
open · @sergqwer · 0 comentarios · En GitHub
Setup & installMulti-GPUNVIDIA / CUDAWindows
Descripción
When the adaptive tier does not wait for its copies (`STRATA_ADAPT_NOWAIT=1`, or `STRATA_ADAPT_LAG` > 1), the CPU can compute an evicted expert from a stale activation. **The bug, as the code runs it:** - `adapt()` marks an evicted expert non-resident in `host_res` at once. - `d_res` follows in `apply_pending`, once the round's copies have landed. Without the wait, the next window runs before that. - `doorbell_publish_res` reads `d_res` and sees the evicted expert as still resident. When every routed expert of a layer looks resident, it does not publish the layer's activation. - The pool plans from `host_res`, so it computes the evicted expert on the CPU, from the activation published for an earlier layer or window. **Measured** with a new check, `STRATA_ADAPT_CHECKRES=1`. Before each window it reads the device tables back, compares them with `host_res`, and counts the layer-windows that computed a CPU expert with no activation published. RTX 5090, IQ2_XS, `--expert-cache 12000`, 32K explain prompt, 600 tokens: | tier mode | windows with mismatched tables | layer-windows with no activation published | |---|---|---| | `STRATA_ADAPT_NOWAIT=1` | 34 of 222 (2,470 entries) | 1 | | `STRATA_ADAPT_NOWAIT=1`, this PR | 0 of 223 | 0 | | `STRATA_ADAPT_LAG=2` | 55 of 223 (3,222 entries: every swap) | 5 | | `STRATA_ADAPT_LAG=2`, this PR | 0 of 218 | 0 | `LAG=2` is reproducible: two runs of each arm give the same tokens. The fix changes its answer from token 228 on. **The change:** - Every verifier launch (`run`, `run_slot_rows`, `batch_launch`, `pl_launch`; a layer split's stages too) runs a pre-launch hook. The hook patches the device tables from a list of the changed entries (`res_patch`, a small kernel on the verifier's stream that reads mapped memory). The whole table is uploaded instead when the list overflows, or when a device does not map it at the host's address. - A failed whole-table upload in `apply_pending` leaves its entries to the hook. A failed patch fails the window instead of launching with stale tables. - It applies only where the tier does not wait. - `--batch-groups` with a layer split publishes every activation instead (`set_always_publish`, which now also turns E-6's device plan off). `--pipeline-windows` keeps its fence. - `STRATA_ADAPT_EVICT_SYNC=0` restores the old behaviour. **Default identity:** with the default (the tier waits) no hook is installed. Tokens and first-token logits are identical to d5ea7133 (p2k, 200 tokens; 32K explain, 600 tokens; `--expert-cache 12000 --pcie-frac 0.25`). **Cost** under `STRATA_ADAPT_NOWAIT=1`, 6,297 slots, 3 interleaved pairs: ms a round -1.1% (chat) and -1.2% (32K explain), i.e. noise. **Not measured:** multi-GPU layer split and peer / helper tiers (one card here; the hook runs per stage and per device, and each device's mapped alias is checked). Found while working on an opt-in that admits the PCIe share's blobs into the VRAM tier: `STRATA_ADAPT_FETCH`, branch `pr-main/tier-fetch-admission`. It builds on this fix and will follow as its own PR. On IQ2_XS with a 6,000-slot cache it measured chat +8% and 32K +4% decode. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
En el sitio
Enlaces a install, modelos, releases.