Issues / #831
#831 The profile's length caps the expert arena — even an explicit --expert-cache N
closed · @ZhongUncle · 2 Kommentare · Auf GitHub
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
With `--expert-profile`, the profile's length becomes the ceiling of the VRAM expert arena: `auto` is capped at it, an explicit `--expert-cache N` above it is silently truncated to it (native pack), and the one path that keeps N slots (`--expert-cache-per-layer`) can never fill the ones past the profile. Measurements below. **Is the profile's length meant to be the ceiling? Or should the profile seed the hot experts and the remaining slots fill on demand** - which is what a user with a trace-built profile (`make_profile.py --no-base`) and a big card most likely expects?
The three mechanisms:
1. `--expert-cache auto` is capped at `profile.size()` (src/program/generate.cpp:3222). Known and documented in tools/make_profile.py (issue #46 was fixed by shipping profiles that rank every pair) - fine.
2. An **explicit** `--expert-cache N` larger than the profile is silently truncated to the profile's length on the native-pack path: the sized-slots loop only iterates the profile's pairs and rewrites `o.expert_cache` (src/program/generate.cpp:3299-3314). No warning is printed.
3. `--expert-cache-per-layer` skips that clamp, so the arena keeps N slots - but the profile is the policy ("the decode-time admission finds no room and every non-profiled expert stays a CPU miss", generate.cpp:3451-3454), so the slots past the profile can never hold anything. The VRAM is allocated and permanently idle.
And since `--serve` requires the profile-filled tier (`Verifier::init`, src/core/verify.cpp:325 - the engine exits at startup without `--expert-profile`), the compulsory-miss fill the engine already implements for profile-less runs (generate.cpp:3445) is unreachable exactly where it would be useful: serving. This all seems adjacent to R4.1's open admission/eviction question (include/strata/core/expert_cache.hpp:107).
## Measured
RTX 3080 20 GB (driver 595.91.07, CUDA 13.2), Xeon E5-2680 v4, 62 GB RAM, Ubuntu 24.04.5, engine 0.1.39 built from source, Qwen3.8-Flash-Next GSQ-RCO IQ2_XS (48 x 512 = 24,576 pairs), `--max-context 262144 --kv int8 --kv-resident 32768 --spec 4 --prefill auto`. Same model and flags throughout; only the profile and `--expert-cache` vary.
**A. Full profile (24,576 pairs), `--expert-cache auto`** - the shipped behavior, fine:
```
strata generate: expert cache 9813 slots, 11.66 GiB of VRAM; policy is
PROFILE, ranked by routing frequency, no eviction.
```
**B. Profile truncated to 3,000 pairs, explicit `--expert-cache 9813`** - asked for 9,813, got 3,000, no warning:
```
strata generate: profile .../expert-profile-3000.bin: 3000 ranked pairs, built for 3000 slots
strata generate: expert cache 3000 slots, 4.04 GiB of VRAM; policy is
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 3000 of 3000 slots from the profile; slot 0 verified
```
**C. Same profile, `--expert-cache-per-layer`, explicit `--expert-cache 8000`** - the arena keeps 8,000 slots (11.25 GiB), 3,000 are filled, the remaining 5,000 (~7 GiB) can never hold an expert:
```
strata generate: expert cache 8000 slots, 11.25 GiB of VRAM; policy is
R4.2g PER-LAYER: each layer owns 166 slots (0..165).
strata generate: pre-filled 3000 of 8000 slots from the profile; slot 0 verified
```
(The 3,000-of-8,000 residency was cross-checked with a local patch that reports `ExpertCache::resident()` in the engine's INFO/DONE lines and the Monitor, where the "experts cached" figure currently reads the capacity instead. Happy to send that as a small PR if you want it.)
## Reproducing
Truncate any profile to its first 3,000 pairs (the format is make_profile.py's):
```python
import sys; sys.path.insert(0, ".")
from tools.make_profile import read_profile, write_profile
write_profile("data/expert-profile-3000.bin", read_profile("data/expert-profile.bin")[:3000])
```
Then start the server with `--expert-profile data/expert-profile-3000.bin` and either `--expert-cache 9813` (case B) or `--expert-cache 8000 --expert-cache-per-layer` (case C), and read the startup log.
## What I'd suggest
- If the ceiling is intended: say so in `--help` / docs/DETAILS.md, and warn when an explicit N is truncated (case B is silent today).
- If not: the smallest step is probably letting the slots past the profile fill on compulsory misses - the admission path exists, it is the verify tier's profile requirement that keeps it out of `--serve`. I can measure the hit-rate difference (static profile vs profile-seeded hybrid) on the 3080 if that helps the R4.1 question.
---
*Drafted with AI assistance (Claude); every number above was measured on the machine listed, and every file:line reference was checked against the source at 0.1.39.*Mehr auf der Site
Links zu Install, Modellen, Releases.