Issues / #1348
#1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)
open · @q8atnight · 1 コメント · GitHub で見る
Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDADocumentationWindows
本文
Hi Niko and everyone, this is a data report and an invitation, not a PR. We spent a day measuring whether experts could be brought into VRAM *before* a window asks for them, instead of the CPU computing the misses. **Short version:** - Misses still cost **up to 34 % of decode time** on our 2× 3090 (95 % hit rate). - That share grows fast as VRAM shrinks: with the cache cut to a 24 GB card's share, the zero-miss headroom is **~+75 %** (bus-limited there). - Two cheap predictors together catch roughly half of the misses, worth an estimated **+12-13 %** here and more on smaller cards. - The catch: the existing adapt swap path is too slow to deliver it, so it needs a low-latency prefetch path in the verify/pool/residency core. Code and tools to reproduce everything are on a public branch (see "Try it / help" at the end). **Setup.** 2× RTX 3090 (x8 Gen4 each, ~13.5 GB/s), Ryzen 3950X, 121 GB DDR4. unsloth UD-Q4_K_XL, v0.1.40.1, `--layer-split 24 --pipeline-windows 2`, locked-RAM arena, 12,779 slots (~52 % of experts), `--spec 4`. After tuning `--pcie-frac 0.2 --spec-min-p 0.7` (+16 % decode vs 0.5/0.5 here; more PCIe share was monotonically *slower*, −17…−32 % at 1.0), real agentic-coding chats run at ~95-110 t/s with a 94-96 % hit rate. **1. The price of a miss is linear and sizeable.** Shrinking the cache via `--vram-reserve-mib` with `STRATA_DECODE_TIMING=1`: | slots | hit | misses per layer-window | ms per window | t/s | |---|---|---|---|---| | 12,779 | 95.3 % | 0.89 | 15.5 | 97 | | 9,712 | 90.3 % | 1.83 | 18.0 | 84 | | 6,363 | 81.0 % | 3.53 | 22.8 | 68 | That's +2.78 ms per window per extra miss per layer-window, i.e. ~0.058 ms per distinct missed expert, with a zero-miss intercept of ~13.0 ms. On real coding windows (wider, T≈2.5) we see ~1.57 misses per layer-window, so "every miss arrives in time" would be worth up to **+34 %** decode on this box. At the 6,363-slot point (≈ the share of experts a single 24 GB card holds) the same arithmetic gives **~+75 %** headroom (22.8 → ~13 ms). There, though, the x8 bus can't carry every miss in time, so a real gain would be well below that ceiling. **2. Misses are predictable, two ways.** Both were measured offline on real sessions: `--dump-routing`, plus a small env-gated dump of each pool call's `x_f` rows (BF16) and ids. Recomputing W_L·x_f reproduces the engine's own top-10 at 99.92 %. The simulator below replays the routing through a mimic of the static profile plus your adapt rules; it matches the engine's hit rate (~94 %) and also reproduces that changing `--adapt-every` / `--adapt-swaps` doesn't move it. - **Recency ring** (cross-window): reserve 16 slots per layer (taken from the coldest residents) that hold experts that just missed, refreshed every window. This gives −24 % distinct misses if a copy lands by the next window. It is very lag-sensitive: −27 / −18 / −12 / −3.5 % at a landing delay of 1 / 2 / 3 / 5 windows. Making adapt itself more aggressive (usage ≥ 1, gain 0-1, every window) only reaches −3…−8 %. It has to be a separate recency tier. - **Router-ahead, zero-shot** (within a window, Fate-style): apply layer L+k's `ffn_gate_inp` to layer L's `x_f`. Token recall of the 10 routed experts: | ahead | @10 | @24 | @48 | |---|---|---|---| | 1 layer | 61 % | 82 % | 90 % | | 2 layers | 52 % | 72 % | 83 % | | 3 layers | 46 % | 66 % | 78 % | Fetching only the best-ranked *non-resident* candidates within what an x8 link moves in 1-3 layers covers **~25-35 % of misses**, mostly different ones from the ring's. - A small trained residual predictor (~200K params per layer) did **not** transfer across sessions (it helped layer 3, hurt layers 10-42), so we'd stay with the router-ahead zero-shot predictor. Combined estimate: ~−45 % misses ≈ **+12-13 % decode** on this box. **3. We built the ring on the adapt path, and it didn't work.** We built the ring inside the serve `adapt()` (an env switch, off by default, 6/6 answers identical with it off). Live result: misses −2…−10 %, speed within noise. Our reading: under `--pipeline-windows 2`, `pl_adapt` → `adapt_fence` → copy → `pl_apply` commit → `exch_ev` release means one ring update per ~4 windows that lands 2-3 windows late. In the serial loop the ring does cut misses −21 %, as simulated, but per-window adapting costs ~2.5 ms per window there, so decode drops 13 %. **What we think a working version needs:** 1. a small kernel in each layer's `pre[l]` graph that applies the next 1-3 routers to `mixed_` rows and publishes top-k with the doorbell (the host is far too slow: ~430 M MAC per window); 2. a per-card prefetch stream into a few reserved slots per layer, with copies issued the moment a layer's routing is known; 3. residency updates that the host plan (`host_res`) and the GPU plan (`d_res`) see identically within a window. This is the part we're least sure how to do safely in your pipeline. **Questions:** 1. Is predictive prefetch something you'd want in Strata (opt-in, byte-identical when off)? Or does it collide with plans you already have for the expert tier? 2. If yes, how would you approach point 3? Is there a safe point per layer where residency can change mid-window, or would you rather double-buffer `d_res` per window? 3. Would the measurement tools help you or others? We have the `x_f` dump switch (~25 lines in `drive_pool_multi`), the trace-replay simulator and the gate scripts, and can open a small PR for the dump switch or share the branch. **Try it / help.** Branch [`q8atnight/Strata:foresight`](https://github.com/q8atnight/Strata/tree/foresight), based on v0.1.40.1 plus our layer-split ports. It has: - the dump switch (`STRATA_FORESIGHT_DUMP`); - the ring reference (`STRATA_FORESIGHT_RING`, measured no gain on the adapt path, kept for comparison); - `tools/foresight/`: router extraction, the trace-replay cache simulator, router-ahead scoring and the Scout trial, with a README. Both switches are off by default and the build is gate-identical. The most useful contributions: - the same three numbers from other setups: price of a miss via `--vram-reserve-mib` + `STRATA_DECODE_TIMING=1`, hit rate, router-ahead coverage. Single 24 GB / 32 GB cards and AMD would tell us where this pays most; - ideas for the residency-consistency part. All numbers here come from one box and one user's real agentic-coding sessions. We're happy to rerun anything on our 2× 3090.
関連リンク
インストール・モデル・リリースへの站内リンク。