Pull requests / #407

#407 Adaptive tier: --adapt-decay, and --adapt-tuned (opt-in): 31% fewer misses, -19% CPU pool, -22% PCIe on UD-Q4_K_XL

closed · @sergqwer · 0 comentários · No GitHub

Setup & installNVIDIA / CUDAModels & quantsDocumentationWindows

Descrição

The adaptive tier swaps the most-routed missing experts into VRAM in place of the least-routed resident ones. It works from routing counts that are multiplied by 0.7 after each pass, every 4 rounds, up to 96 swaps.

This PR adds:
- **`--adapt-decay F`**: sets that factor (default 0.7, unchanged).
- **`--adapt-tuned`** (`STRATA_ADAPT_TUNED=1`), **off by default**:
  - It re-ranks every 2 rounds, up to 192 swaps, with counts x0.92: a longer memory, and more frequent steps.
  - It applies only when the cache holds 20-60% of the experts and every expert outside VRAM is in RAM. With a RAM budget, that means the budget holds all of them; otherwise a swap can read the drive.
  - Outside those conditions the defaults stay. Explicit `--adapt-*` flags win.
  - The start log names the choice: `--adapt-tuned: the cache holds 30% of the experts -> adaptive tier every 2 rounds, 192 swaps, x0.92`.

It is opt-in because more experts on the GPU change which kernels compute them, so it changes tokens and would not pass a byte-identical gate.

## How it was chosen

We recorded two decode routing traces (`--dump-routing`) and replayed them through the tier's policy. The replay matched the engine's misses within 3%.
- **A window of the last 8-16 tokens was worse:** 24-96% more misses. It evicts experts a conversation uses rarely but steadily, and swaps them back and forth.
- **Longer memory with more frequent steps was better:** 2 / 192 / x0.92 gave 31-38% fewer misses.
- **An LRU tail** cost 8-9x the copies.

## Measured on Unsloth UD-Q4_K_XL

Setup:
- RTX 5090, 9950X3D, 128 GB, Windows 11;
- the import from docs/UNSLOTH_Q4.md, with `--resident-budget-gib 90` (every expert outside VRAM in RAM), `--max-context 32768 --kv int8 --vram-reserve-mib 1500`;
- cache auto: 7,542 slots = 30% of the experts;
- a 1,000-token chat; the runs alternated between the defaults and `--adapt-every 2 --adapt-swaps 192 --adapt-decay 0.92`, the same as `--adapt-tuned` here.

| 6 pairs | defaults | tuned |
| --- | ---: | ---: |
| misses per layer (CPU + PCIe share) | 4.06 | **2.79 (-31%)** |
| hit rate | 0.908 | **0.934** |
| CPU expert pool, ms a round | 9.9 | **8.0 (-19%)** |
| PCIe traffic a round (PCIe share + swaps) | 329 MB | **256 MB (-22%)** |
| swaps a round | 20 | 28 |
| round | 32.0 +- 3.0 ms | 32.3 +- 2.2 ms |

**With 16 busy CPU worker processes** beside the engine (a desktop doing other work), 4 pairs:
- misses -18%;
- CPU pool 28.8 -> 23.1 ms a round;
- round 53.9 +- 1.9 -> 52.2 +- 5.9 ms.

**What the numbers say:**
- **Steady wins:** fewer misses, a lighter CPU pool and less PCIe traffic.
- **The round time does not move measurably here.**
  - In most layers the GPU's work, not the CPU's, sets a layer's time. So only the smaller PCIe share shortens the GPU side: ~95 MB a round, ~2 ms of 32.
  - That is below the run-to-run spread, because every greedy run takes its own token path.
- **Other configurations:**
  - Our NVFP4 fork, same GPU, 34% coverage, 6 pairs: 19.9 -> 17.8 ms a round.
  - IQ2_XS showed no gain at 14%, 38%, 49% or 62% coverage, which is why the rule stays inside 20-60% with the experts in RAM.
- **One cost:** in the RAM-budget mode the adaptive thread's host time goes from ~1.5 to ~3.1 ms a round. It runs beside the commit.

Ideas we did not get to:
- The swaps land in bursts of up to 192 copies (~0.5 GB) every 2 rounds, while the PCIe share's fetch sits on the critical path. Spreading them, or scheduling them where the bus is idle, might turn the smaller traffic into round time.
- A deterministic A/B: verify windows fed a fixed token file, to resolve 1-2 ms differences. It would help any decode tuning.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

No site

Links install, modelos, releases.