Issues / #1348

#1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)

open · @q8atnight · 1 评论 · 在 GitHub 查看

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDADocumentationWindows

描述

Hi Niko and everyone, this is a data report and an invitation, not a PR. We spent a day measuring whether experts
could be brought into VRAM *before* a window asks for them, instead of the CPU computing the misses.

**Short version:**
- Misses still cost **up to 34 % of decode time** on our 2× 3090 (95 % hit rate).
- That share grows fast as VRAM shrinks: with the cache cut to a 24 GB card's share, the zero-miss headroom is **~+75 %**
  (bus-limited there).
- Two cheap predictors together catch roughly half of the misses, worth an estimated **+12-13 %** here and more on
  smaller cards.
- The catch: the existing adapt swap path is too slow to deliver it, so it needs a low-latency prefetch path in the
  verify/pool/residency core.

Code and tools to reproduce everything are on a public branch (see "Try it / help" at the end).

**Setup.** 2× RTX 3090 (x8 Gen4 each, ~13.5 GB/s), Ryzen 3950X, 121 GB DDR4. unsloth UD-Q4_K_XL, v0.1.40.1,
`--layer-split 24 --pipeline-windows 2`, locked-RAM arena, 12,779 slots (~52 % of experts), `--spec 4`.
After tuning `--pcie-frac 0.2 --spec-min-p 0.7` (+16 % decode vs 0.5/0.5 here; more PCIe share was monotonically
*slower*, −17…−32 % at 1.0), real agentic-coding chats run at ~95-110 t/s with a 94-96 % hit rate.

**1. The price of a miss is linear and sizeable.** Shrinking the cache via `--vram-reserve-mib` with
`STRATA_DECODE_TIMING=1`:

| slots | hit | misses per layer-window | ms per window | t/s |
|---|---|---|---|---|
| 12,779 | 95.3 % | 0.89 | 15.5 | 97 |
| 9,712 | 90.3 % | 1.83 | 18.0 | 84 |
| 6,363 | 81.0 % | 3.53 | 22.8 | 68 |

That's +2.78 ms per window per extra miss per layer-window, i.e. ~0.058 ms per distinct missed expert, with a
zero-miss intercept of ~13.0 ms. On real coding windows (wider, T≈2.5) we see ~1.57 misses per layer-window, so
"every miss arrives in time" would be worth up to **+34 %** decode on this box. At the 6,363-slot point (≈ the share
of experts a single 24 GB card holds) the same arithmetic gives **~+75 %** headroom (22.8 → ~13 ms). There, though,
the x8 bus can't carry every miss in time, so a real gain would be well below that ceiling.

**2. Misses are predictable, two ways.** Both were measured offline on real sessions: `--dump-routing`, plus a small
env-gated dump of each pool call's `x_f` rows (BF16) and ids. Recomputing W_L·x_f reproduces the engine's own top-10
at 99.92 %. The simulator below replays the routing through a mimic of the static profile plus your adapt rules; it
matches the engine's hit rate (~94 %) and also reproduces that changing `--adapt-every` / `--adapt-swaps` doesn't
move it.

- **Recency ring** (cross-window): reserve 16 slots per layer (taken from the coldest residents) that hold experts
  that just missed, refreshed every window. This gives −24 % distinct misses if a copy lands by the next window. It is
  very lag-sensitive: −27 / −18 / −12 / −3.5 % at a landing delay of 1 / 2 / 3 / 5 windows. Making adapt itself more
  aggressive (usage ≥ 1, gain 0-1, every window) only reaches −3…−8 %. It has to be a separate recency tier.
- **Router-ahead, zero-shot** (within a window, Fate-style): apply layer L+k's `ffn_gate_inp` to layer L's `x_f`.
  Token recall of the 10 routed experts:

  | ahead | @10 | @24 | @48 |
  |---|---|---|---|
  | 1 layer | 61 % | 82 % | 90 % |
  | 2 layers | 52 % | 72 % | 83 % |
  | 3 layers | 46 % | 66 % | 78 % |

  Fetching only the best-ranked *non-resident* candidates within what an x8 link moves in 1-3 layers covers
  **~25-35 % of misses**, mostly different ones from the ring's.
- A small trained residual predictor (~200K params per layer) did **not** transfer across sessions (it helped layer 3,
  hurt layers 10-42), so we'd stay with the router-ahead zero-shot predictor.

Combined estimate: ~−45 % misses ≈ **+12-13 % decode** on this box.

**3. We built the ring on the adapt path, and it didn't work.** We built the ring inside the serve `adapt()` (an env
switch, off by default, 6/6 answers identical with it off). Live result: misses −2…−10 %, speed within noise. Our
reading: under `--pipeline-windows 2`, `pl_adapt` → `adapt_fence` → copy → `pl_apply` commit → `exch_ev` release
means one ring update per ~4 windows that lands 2-3 windows late. In the serial loop the ring does cut misses −21 %,
as simulated, but per-window adapting costs ~2.5 ms per window there, so decode drops 13 %.

**What we think a working version needs:**
1. a small kernel in each layer's `pre[l]` graph that applies the next 1-3 routers to `mixed_` rows and publishes
   top-k with the doorbell (the host is far too slow: ~430 M MAC per window);
2. a per-card prefetch stream into a few reserved slots per layer, with copies issued the moment a layer's routing is
   known;
3. residency updates that the host plan (`host_res`) and the GPU plan (`d_res`) see identically within a window.
   This is the part we're least sure how to do safely in your pipeline.

**Questions:**
1. Is predictive prefetch something you'd want in Strata (opt-in, byte-identical when off)? Or does it collide with
   plans you already have for the expert tier?
2. If yes, how would you approach point 3? Is there a safe point per layer where residency can change mid-window, or
   would you rather double-buffer `d_res` per window?
3. Would the measurement tools help you or others? We have the `x_f` dump switch (~25 lines in `drive_pool_multi`),
   the trace-replay simulator and the gate scripts, and can open a small PR for the dump switch or share the branch.

**Try it / help.** Branch [`q8atnight/Strata:foresight`](https://github.com/q8atnight/Strata/tree/foresight), based on
v0.1.40.1 plus our layer-split ports. It has:
- the dump switch (`STRATA_FORESIGHT_DUMP`);
- the ring reference (`STRATA_FORESIGHT_RING`, measured no gain on the adapt path, kept for comparison);
- `tools/foresight/`: router extraction, the trace-replay cache simulator, router-ahead scoring and the Scout trial,
  with a README.

Both switches are off by default and the build is gate-identical. The most useful contributions:
- the same three numbers from other setups: price of a miss via `--vram-reserve-mib` + `STRATA_DECODE_TIMING=1`, hit
  rate, router-ahead coverage. Single 24 GB / 32 GB cards and AMD would tell us where this pays most;
- ideas for the residency-consistency part.

All numbers here come from one box and one user's real agentic-coding sessions. We're happy to rerun anything on our
2× 3090.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。