Pull requests / #1668

#1668 Fewer missed experts in decode: `STRATA_ROUTE_TAIL_SKIP=R` and `STRATA_ROUTE_PRIOR=lambda` (decode x1.33 on a 12 GB card, opt-in, change the output)

open · @LeGeRyChEeSe · 0 comments · View on GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows

Description

## Title
Issue: related: `STRATA_ROUTE_RESIDENT` (5df35dc, experimental) attacks the same misses by another rule; the numbers
below compare all three. The prior follows "Mixture of Cache-Conditional Experts" (arXiv 2412.00099), used here
without any training. No issue or PR found for either rule (searched 2026-10-09: tail, rank, skip, prior,
cache-conditional, missed experts).

## Summary
**Why it matters:** on a card that cannot hold every expert, the decode time that is not GPU compute goes to the
missed experts: their copy over PCIe and the CPU pool's rows. Two opt-in switches cut those misses, each on its own
variable. Both are off by default, and the default path is unchanged.

- **`STRATA_ROUTE_TAIL_SKIP=R`**: a missed expert (not in VRAM, nor on a helper or peer GPU) that **every** token of
  the verify window routes at rank R or lower in weight is neither copied nor computed. Its contribution is zero, and
  the other experts keep the router's weights (no renormalisation). An expert that any token ranks above R is served
  as before, so the window's strongest choices are never dropped.
- **`STRATA_ROUTE_PRIOR=lambda`**: each token's top-10 is picked again with `lambda` added to the logit of every
  expert resident in this card's cache. The weights stay the router's probabilities (softmax over the 512 unbiased
  logits), renormalised over the 10 picked. Unlike `STRATA_ROUTE_RESIDENT`, which swaps a missed rank-6-9 expert for
  the best resident one within a margin, every rank can move, by an amount the logits decide.

RTX 4070 SUPER 12 GB, IQ2_XS, 23 coding-agent turns (2.5K-103K tokens), median of the per-turn decode ratios:

| | to the default | turns faster | expert cache hit rate |
|---|---|---|---|
| default (63.5 tok/s median) | - | - | 0.75 |
| `STRATA_ROUTE_TAIL_SKIP=7` | x1.17 (x1.03 to x1.39) | 23/23 | 0.81 |
| **`STRATA_ROUTE_TAIL_SKIP=7` + `STRATA_ROUTE_PRIOR=0.5`** | **x1.33** (x1.13 to x1.72) | **23/23** | 0.90 |

Measured again on top of HC-REQ8 - a local requantization of the hyper-connection projections at load, the approach
of #1262 (not part of this PR), which gives this card a 314-slot larger expert cache - against that base:

| on top of HC-REQ8 | to that base | turns faster |
|---|---|---|
| `STRATA_ROUTE_TAIL_SKIP=7` | x1.165 (x0.93 to x1.40) | 20/23 |
| `STRATA_ROUTE_PRIOR=0.5` | x1.162 (x1.04 to x1.32) | 23/23 |
| both | x1.354 (x1.16 to x1.84) | 23/23 |
| `STRATA_ROUTE_RESIDENT=1`, for comparison | x1.163 (x1.02 to x1.37) | 23/23 |
| both, cache forced to 2000 slots (a smaller card), to the default | x1.41 (x1.28 to x1.66) | 23/23 |

With HC-REQ8 and both switches, the median decode of these turns goes from 64.0 to 90.5 tok/s against the
default (x1.38, 23/23 faster). Prompt reads are unchanged: both rules act only in the decode windows.

## What changed
- `src/core/expert_source.cpp`, `include/strata/core/expert_source.hpp` (tail skip):
  - in `expert_pool_dispatch_multi`, after the distinct experts are counted, the skipped ones get `kind = -2`;
  - their rows are zeroed, and they take no PCIe slot and no CPU job;
  - `any_cpu` and the prefetch list count only `kind == -1`;
  - the rule is not applied with a peer or helper GPU;
  - two counters are added to `ExpertDispatch`.
- `src/program/generate.cpp`: one cumulative log line per request,
  `route tail skip: N missed experts skipped (M entries) since the start`.
- `src/kernels/cuda/route_prior.cu`, `include/strata/kernels/route_prior.hpp` (new): one warp per token, 16 logits
  per lane, the same warp idiom as `native_router.cu`.
- `src/core/verify.cpp`: the prior is applied after the router and before `STRATA_ROUTE_RESIDENT` and the plan, when
  the variable is set, NE == 512 and K == 10.
- `CMakeLists.txt`: the new kernel is added to `strata_kernels`.
- `tools/agent_turns_bench.py` (new): the workload below as a reusable A/B, in the manner of `ab_engine.py`; it
  prints the median of the per-turn decode ratios and the turns each arm won.

## Extra Notes
- **Measured on:** engine 0.1.41 (`fb58e0d`) plus this change. RTX 4070 SUPER 12 GB (PCIe 4.0 x16), Ryzen 5 5600X,
  64 GB DDR4, Windows 11, driver 617.14, CUDA 13.1, IQ2_XS native pack, `--spec 4 --mtp`, `--kv int8
  --kv-resident 32768`, `--vram-reserve-mib 1000`.
- **Method:** the same binary in every arm, only the environment differs. Whole engine runs, interleaved and
  mirrored (for example A B C D D C B A). T = 1, top-p 0.95, 256 new tokens per turn.
- **Noise:** the default against itself gave x1.005 (x0.88 to x1.15) and x1.023 (x0.77 to x1.09) in the two sessions.
- **Workload:** `tools/agent_turns_bench.py`, added by this PR: five simulated coding-agent sessions over this
  repository's own files (system prompt, four tools, real files as tool results, at most 2000 lines per read), 23
  turns from ~2.5K to ~100K tokens. `python tools/agent_turns_bench.py strata-iq2_xs.json base: skip:STRATA_ROUTE_TAIL_SKIP=7`.
- **Quality**, stated plainly: both rules change the output. Teacher-forced KL against the plain default over 600
  positions on 3 public prompts (6K / 18K / 27K tokens), `STRATA_LOGPOS`, adaptive tier frozen:

| | KL mean | perplexity (default: 3.92 / 34.08 / 30.44) |
|---|---|---|
| control: default with a smaller cache (3300 slots, placement noise) | 0.21 / 0.93 / 0.75 | 3.87 / 34.10 / 28.68 |
| `STRATA_ROUTE_TAIL_SKIP=7` | 0.27 / 1.24 / 0.90 | 3.59 / 32.16 / 29.87 |
| `STRATA_ROUTE_TAIL_SKIP=7` + `STRATA_ROUTE_PRIOR=0.5` | 0.38 / 1.36 / 1.11 | 3.85 / 36.39 / 32.01 |
| on top of HC-REQ8: + tail skip | 0.31 / 1.27 / 0.94 | 3.66 / 32.37 / 27.21 |
| on top of HC-REQ8: + prior | 0.35 / 1.36 / 1.05 | 4.29 / 37.50 / 28.81 |
| on top of HC-REQ8: + both | 0.42 / 1.49 / 1.12 | 4.05 / 37.90 / 31.15 |
| on top of HC-REQ8: + `STRATA_ROUTE_RESIDENT=1` | 0.38 / 1.46 / 1.06 | 3.90 / 34.19 / 28.48 |

The tail skip moves the output least, for the same speed as `STRATA_ROUTE_RESIDENT=1`. The prior adds speed and
moves it more.

HumanEval (164, greedy, thinking off) does not separate the arms:

| arm | passed |
|---|---|
| default, two runs | 157 and 156 |
| default with a smaller cache | 154 |
| on top of HC-REQ8: tail skip | 154 |
| on top of HC-REQ8: prior | 153 |
| on top of HC-REQ8: both | 154 |
| on top of HC-REQ8: `STRATA_ROUTE_RESIDENT=1` | 155 |

The problems that flip are the same handful (HumanEval/75, 108, 130, 132, 140, 147) in every arm, the controls
included. Greedy decoding on this engine is not repeatable across runs once the expert placement changes (GPU and
CPU experts round differently), so at n = 164 the KL is the finer measure.

- **On the values:** R = 8 gave x1.09 on 0.1.39 and R = 7 x1.16. Lambda 0.75 with R = 6 gained a few percent more
  on 0.1.39, with a larger KL. 7 and 0.5 are the values measured here.
- **Not checked:**
  - HIP and SYCL: CUDA build only. The skip is host code shared with HIP; the kernel uses the warp idiom of
    `native_router.cu`; the SYCL tree is untouched.
  - a layer split and helper GPUs: the skip is off with them;
  - the Coder: the prior needs NE == 512 and K == 10 and stays off otherwise; the skip was not measured there.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.