Pull requests / #1089
#1089 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing traces
open · @ZhongUncle · 0 Kommentare · Auf GitHub
Multi-GPUNVIDIA / CUDAModels & quantsWindows
Beschreibung
Repost of #991, which GitHub closed automatically when the repository's history was cleaned up — not a rejection, per the maintainer's note there. Same content, rebuilt on the new `main`.
`ExpertCache::admit` "never evicts, because eviction policy is a measured question (`R4.1`'s LFU-decay vs LRU sweep) and a placeholder policy would set the hit rate that everything downstream is then sized against" (`expert_cache.hpp:133-135`, repeated at `expert_source.hpp:259-261`). This PR is that measurement. It adds the harness and reports the sweep on real routing traces — it deliberately does **not** implement an eviction policy; what Phase 3 puts in place of the current object (`expert_source.hpp:411`) stays the maintainer's call, now with numbers.
**TL;DR**
- Eviction is worth 48.6–55.5 hit-rate points in the 2000–8000 slot band (LRU 79.9% vs compulsory-miss 24.3% @ 4105 slots), still +42.2 at 10377
- LRU stays within 1.8–11 points of the oracle at 4105 slots and above on the session and conversations (individual single-turn tasks: up to 14); a tuned LFU-decay matches it there only with a workload-dependent window (at 2000 slots even tuned trails by 3.5–4.9 points), and the untuned default (w=65536) loses to it from 2000 slots up
- Details in [Findings](#findings); limits in [Limits](#limits)
### What it adds
Two new files, no engine changes:
- `tools/bench_eviction.py` — replays `--dump-routing` traces through simulated slot pools. A policy's hit rate depends only on the order `(layer, expert)` pairs were routed, which the trace records, so the replay needs no GPU. Policies: **compulsory-miss** (the engine's current behaviour; `fill-only` in the CLI), **lru**, **lfu-decay** (tunable window), **oracle** (Belady's MIN, the upper bound). Global scope or **R4.2g per-layer admission** matching `ExpertCache::layer_slot_range`; any slot count; optional `--expert-profile` prefill; per-window rates so a topic shift shows up as a dip.
- `tools/test_bench_eviction.py` — 14 cases, no GPU: each policy's victim choice, the per-layer quota arithmetic, the windows split, profile prefill order.
### Measured results
Traces collected on:
- RTX 3080 20 GB (driver 595.91.07, CUDA 13.2)
- Xeon E5-2680 v4
- 62 GB RAM
- Ubuntu 24.04
- Qwen3.8-Flash-Next IQ2_XS pack (48 layers × 512 experts = 24,576 possible pairs)
- `--spec 4 --spec-min-p 0.5`
All traces are **decode-phase** (`--dump-routing` writes no prefill records; with `--spec` on, a trace position is a verify-window candidate, rejected drafts included — the true expert load).
Workloads:
- 8 single-turn tasks (~2.1–2.4M lookups each)
- a 7-task session — all of the above except the ancient-japan single-topic control, replayed in order (16,025,760 lookups, 24,533 distinct pairs)
- two multi-turn conversations (10 and 12 turns)
**Session hit rate by slots (cold start), in % (100 = every lookup resident):**
| slots | compulsory-miss | lru | lfu-decay (w=65536) | lfu-decay (best w) | best w | oracle |
|---|---|---|---|---|---|---|
| 2000 | 10.88 | 61.92 | 44.42 | 58.37 | 16384 | 80.21 |
| 4105 | 24.32 | 79.86 | 67.23 | 80.56 | 16384 | 90.63 |
| 8000 | 44.44 | 93.04 | 89.56 | 93.36 | 16384 | 96.97 |
| 10377 | 53.89 | 96.11 | 95.54 | 96.23 | 32768 | 98.33 |
(10377 ≈ what `--expert-cache auto` picks on the 20 GB card.)
**LFU-decay window sweep (session, cold):**
| window | 2000 | 4105 | 8000 | 10377 |
|---|---|---|---|---|
| 4096 | 53.50 | 66.55 | 77.87 | 84.84 |
| 8192 | 56.18 | 71.38 | 82.70 | 89.22 |
| 16384 | 58.37 | 80.56 | 93.36 | 96.20 |
| 32768 | 51.66 | 74.72 | 93.16 | 96.23 |
| 65536 | 44.42 | 67.23 | 89.56 | 95.54 |
| lru (reference) | 61.92 | 79.86 | 93.04 | 96.11 |
**R4.2g per-layer scope (slots split evenly across 48 layers, session, cold):**
| slots | compulsory-miss | lru | lfu-decay (w=65536) | oracle |
|---|---|---|---|---|
| 384 (8/layer) | 1.20 | 10.30 | 10.72 | 42.39 |
| 2000 (~42/layer) | 11.06 | 61.84 | 43.46 | 78.63 |
| 3072 (64/layer) | 18.44 | 72.23 | 56.89 | 85.64 |
| 4105 (~86/layer) | 24.26 | 78.86 | 66.85 | 89.65 |
| 8000 (~167/layer) | 44.02 | 91.85 | 88.35 | 96.39 |
| 10377 (~216/layer) | 53.76 | 95.32 | 94.26 | 97.95 |
The two reference points quoted in `expert_cache.hpp:153-154` are consistent with these traces:
- compulsory-miss at 8 slots/layer is 1.2% — the same pathology the comment reports as **2.97%** @ 256 slots global
- LRU at 64 slots/layer (72.2%) lands next to the **70.4%** the comment quotes from `Memory/cache_allocation.py`
### Findings
1. **Eviction is worth 48.6–55.5 hit-rate points in the 2000–8000 slot band** (+51.0 @ 2000, +55.5 @ 4105, +48.6 @ 8000), still +42.2 at 10377. LRU stays within 1.8–11 points of the oracle at 4105 slots and above on the session and conversations — individual single-turn tasks run higher, up to 14.1 (data-analysis) and 13.4 (devops-troubleshooting), so LRU does not close the gap everywhere. At 2000 slots the gap widens to 16–18 points (session 18.3, conv1 15.8, conv2 16.6, per-layer 16.8), so caches that small leave headroom a smarter policy could claim — though 2000 is already below what `--expert-cache auto` picks on this card.
2. **Untuned LFU-decay (w=65536, the upstream default) loses to LRU at every slot count from 2000 up.** The one exception is the tiny 384-slot per-layer cell (lfu-decay 10.72% vs lru 10.30%) — at 8 slots per layer neither policy has room to work. The default window is ~4× wrong for these workloads.
3. **A tuned LFU-decay matches but does not robustly beat LRU.** With the right window it edges ahead on some workloads (japan-travel @ 4105: 82.3% at w=8192 vs LRU 80.7%; session @ 4105: 80.6% at w=16384 vs 79.9%), but the winning window moves with the workload (8192 on japan-travel, 16384 on the session — sweep table in the appendix), and at 10377 slots w=16384 and w=32768 tie within 0.03 points. At 2000 slots even the best window trails LRU by 3.5–4.9 points (global 58.4% vs 61.9%; per-layer 56.9% vs 61.8%). No fixed window dominates LRU — that is the measured answer to R4.1's sweep question for these workloads.
4. **Multi-turn reuse lifts compulsory-miss substantially**: 43.0% @ 4105 in the 10-turn single-topic conversation vs 24.3% on the single-turn session (miss rate 57.0% vs 75.7%; note this compares two workload types — consecutive turns vs separate tasks). The penalty returns on topic pivots: the travel→code→math conversation drops to 30.1%.
5. **Profile prefill mainly helps compulsory-miss** (+6 points @ 4105), ~0 for LRU. A warm start, not a substitute for eviction.
6. Conclusions are **stable across scopes** (global vs per-layer): the untuned default loses to LRU in both (from 2000 slots up — the 384-slot exception is in finding 2), the tuned best matches-but-does-not-robustly-beat LRU in both, and both decay sweeps peak at w=16384 in the 2000–8000 band (sweep table in the appendix; at 10377 the top is flat, see finding 3).
### Design checks
All pass:
- coverage 24,533 > largest slot count
- worst pairwise top-500 Jaccard 0.138
- oracle − compulsory-miss = 66 points at 4105 slots
- fill-only's per-window rate drops 34% → 19% across the session (from the `--windows 4` per-window rates; the routing shift is real)
### Validation against the engine's own counters
- **What was compared:** the replay in `tools/bench_eviction.py` ("the sim" below — it is this PR's tool, the engine is untouched) was reconciled against `--stats` on matched runs. The engine counts per-token entries — the sim's basis — but its reported lookups (hits + refused) exclude the misses it serves over PCIe
- **The PCIe rule** (`expert_source.cpp:2030-2209`): distinct misses ranked by first occurrence per window-layer, the last `m = (nmiss × 79) >> 8` fetched via PCIe as `offload_entries`; 79/256 = this host's measured pcie_frac 0.31
- **Result:** replaying with that one rule, the sim reproduces the engine's lookups within **0.2%** and refused within **0.75%**; the residual on the hit count (+2.0% @ 2000 slots, +3.2% @ 4105 — relative to the hit count, not the hit-rate points in Limits) is the prefill-phase admissions the sim deliberately does not model — under 1 hit-rate point once the PCIe rule is applied
- **Real-config run** (profile + the adaptive tier, 4317 adaptive swaps, 10508 slots): the engine reports h = 0.9729, between the sim's compulsory-miss (0.3064) and LRU (0.9925) at the same slot count — consistent with a tier that swaps but does not behave like a full LRU (an interpretation of one number, not a proof). (These three numbers are on that run's own trace, profile-prefilled — not the session trace in the tables above.)
All sim policies share one accounting basis, so the policy comparisons are unaffected by these differences.
### Limits
- Traces cover the **decode phase** only (`--dump-routing` writes nothing during prefill)
- The sim models a bare `ExpertCache` with one policy: no adaptive tier, no PCIe offload, no pack-size variation. Measured on the two static arms, the engine's h reads **2.9 and 4.8 hit-rate points** higher than the raw sim rate (0.1333 vs 0.1044 @ 2000 slots; 0.2361 vs 0.1878 @ 4105); most of that is the PCIe-offload denominator, and under 1 point remains once that rule is applied
- One model, one quant, one GPU. The workloads are synthetic prompts on real routing
- A hit rate is necessary but not sufficient — the kernel side (R4.2c on) is untouched
### Reproduce
```bash
# collect (host with the model; one run per prompt file, pretokenized via the pack's chat template)
./engine/strata --pack Strata-data/packs/iq2_xs --native <shard1.gguf> --ple-gguf <shard2.gguf> \
--expert-profile data/expert-profile.bin --expert-cache auto --prefill auto \
--spec 4 --spec-min-p 0.5 --mtp Strata-data/mtp/rt \
--max-context 24576 --kv int8 --kv-resident 32768 \
--tokens-file PROMPT.tokens --max-new 4096 --dump-routing task.trace.bin
# replay (no GPU)
python3 tools/bench_eviction.py task1.trace.bin task2.trace.bin ... \
--slots 2000,4105,8000,10377 --windows 4 --csv session.csv
python3 tools/bench_eviction.py session.trace.bin --slots 2000,4105 --decay-window 16384
python3 tools/bench_eviction.py session.trace.bin --slots 4105 --per-layer
python3 tools/test_bench_eviction.py
```
<details>
<summary>Per-task, conversation, and single-task sweep tables</summary>
japan-travel single-task LFU-decay window sweep @ 4105 slots, cold (supports finding 3; lru for reference):
| window | 2048 | 4096 | 8192 | 16384 | 32768 | 65536 | lru |
|---|---|---|---|---|---|---|---|
| hit rate | 79.01 | 81.50 | **82.30** | 81.92 | 76.87 | 70.95 | 80.72 |
Per-layer LFU-decay window sweep, session, cold (supports finding 6; lru for reference):
| window | 4096 | 8192 | 16384 | 32768 | 65536 | lru |
|---|---|---|---|---|---|---|
| 2000 slots | 52.08 | 54.64 | **56.90** | 50.62 | 43.46 | 61.84 |
| 4105 slots | 65.06 | 70.10 | **79.32** | 73.87 | 66.85 | 78.86 |
| 8000 slots | 77.73 | 82.54 | **92.39** | 91.86 | 88.35 | 91.85 |
Per-task @ 4105 slots, cold:
| task | compulsory-miss | lru | lfu-decay | oracle |
|---|---|---|---|---|
| japan-travel | 44.21 | 80.72 | 70.95 | 90.97 |
| cross-language-code | 28.08 | 82.05 | 70.36 | 91.33 |
| math-learning | 26.34 | 81.40 | 67.65 | 90.69 |
| ancient-japan (control, single-topic) | 47.37 | 84.01 | 74.71 | 92.26 |
| chinese-report | 36.44 | 80.24 | 66.00 | 90.51 |
| data-analysis | 14.85 | 74.97 | 68.06 | 89.09 |
| creative-writing | 23.30 | 82.92 | 70.85 | 91.85 |
| devops-troubleshooting | 24.94 | 75.59 | 70.17 | 88.94 |
conv1 (10-turn, single topic) / conv2 (12-turn, travel→code→math), cold:
| slots | conv1 compulsory-miss | conv1 lru | conv1 oracle | conv2 compulsory-miss | conv2 lru | conv2 oracle |
|---|---|---|---|---|---|---|
| 2000 | 12.82 | 67.28 | 83.07 | 13.11 | 64.83 | 81.47 |
| 4105 | 43.02 | 83.69 | 92.26 | 30.06 | 81.71 | 91.21 |
| 8000 | 72.82 | 94.17 | 97.33 | 49.72 | 93.15 | 96.96 |
| 10377 | 82.08 | 96.64 | 98.46 | 60.47 | 96.14 | 98.27 |
Session @ 4105 slots with `--expert-pMehr auf der Site
Links zu Install, Modellen, Releases.