Pull requests / #407
#407 Adaptive tier: --adapt-decay, and --adapt-tuned (opt-in): 31% fewer misses, -19% CPU pool, -22% PCIe on UD-Q4_K_XL
closed · @sergqwer · 0 コメント · GitHub で見る
Setup & installNVIDIA / CUDAModels & quantsDocumentationWindows
本文
The adaptive tier swaps the most-routed missing experts into VRAM in place of the least-routed resident ones. It works from routing counts that are multiplied by 0.7 after each pass, every 4 rounds, up to 96 swaps. This PR adds: - **`--adapt-decay F`**: sets that factor (default 0.7, unchanged). - **`--adapt-tuned`** (`STRATA_ADAPT_TUNED=1`), **off by default**: - It re-ranks every 2 rounds, up to 192 swaps, with counts x0.92: a longer memory, and more frequent steps. - It applies only when the cache holds 20-60% of the experts and every expert outside VRAM is in RAM. With a RAM budget, that means the budget holds all of them; otherwise a swap can read the drive. - Outside those conditions the defaults stay. Explicit `--adapt-*` flags win. - The start log names the choice: `--adapt-tuned: the cache holds 30% of the experts -> adaptive tier every 2 rounds, 192 swaps, x0.92`. It is opt-in because more experts on the GPU change which kernels compute them, so it changes tokens and would not pass a byte-identical gate. ## How it was chosen We recorded two decode routing traces (`--dump-routing`) and replayed them through the tier's policy. The replay matched the engine's misses within 3%. - **A window of the last 8-16 tokens was worse:** 24-96% more misses. It evicts experts a conversation uses rarely but steadily, and swaps them back and forth. - **Longer memory with more frequent steps was better:** 2 / 192 / x0.92 gave 31-38% fewer misses. - **An LRU tail** cost 8-9x the copies. ## Measured on Unsloth UD-Q4_K_XL Setup: - RTX 5090, 9950X3D, 128 GB, Windows 11; - the import from docs/UNSLOTH_Q4.md, with `--resident-budget-gib 90` (every expert outside VRAM in RAM), `--max-context 32768 --kv int8 --vram-reserve-mib 1500`; - cache auto: 7,542 slots = 30% of the experts; - a 1,000-token chat; the runs alternated between the defaults and `--adapt-every 2 --adapt-swaps 192 --adapt-decay 0.92`, the same as `--adapt-tuned` here. | 6 pairs | defaults | tuned | | --- | ---: | ---: | | misses per layer (CPU + PCIe share) | 4.06 | **2.79 (-31%)** | | hit rate | 0.908 | **0.934** | | CPU expert pool, ms a round | 9.9 | **8.0 (-19%)** | | PCIe traffic a round (PCIe share + swaps) | 329 MB | **256 MB (-22%)** | | swaps a round | 20 | 28 | | round | 32.0 +- 3.0 ms | 32.3 +- 2.2 ms | **With 16 busy CPU worker processes** beside the engine (a desktop doing other work), 4 pairs: - misses -18%; - CPU pool 28.8 -> 23.1 ms a round; - round 53.9 +- 1.9 -> 52.2 +- 5.9 ms. **What the numbers say:** - **Steady wins:** fewer misses, a lighter CPU pool and less PCIe traffic. - **The round time does not move measurably here.** - In most layers the GPU's work, not the CPU's, sets a layer's time. So only the smaller PCIe share shortens the GPU side: ~95 MB a round, ~2 ms of 32. - That is below the run-to-run spread, because every greedy run takes its own token path. - **Other configurations:** - Our NVFP4 fork, same GPU, 34% coverage, 6 pairs: 19.9 -> 17.8 ms a round. - IQ2_XS showed no gain at 14%, 38%, 49% or 62% coverage, which is why the rule stays inside 20-60% with the experts in RAM. - **One cost:** in the RAM-budget mode the adaptive thread's host time goes from ~1.5 to ~3.1 ms a round. It runs beside the commit. Ideas we did not get to: - The swaps land in bursts of up to 192 copies (~0.5 GB) every 2 rounds, while the PCIe share's fetch sits on the critical path. Spreading them, or scheduling them where the bus is idle, might turn the smaller traffic into round time. - A deterministic A/B: verify windows fed a fixed token file, to resolve 1-2 ms differences. It would help any decode tuning. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio
関連リンク
インストール・モデル・リリースへの站内リンク。