Pull requests / #796
#796 generate: one startup VRAM plan before the expert cache is committed (#765)
closed · @j-luwierski · 0 Kommentare · Auf GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindowsLinux
Beschreibung
Closes #765
## Problem
At very large contexts, individually reasonable VRAM decisions combined into an
unworkable whole. Startup sized the expert cache from `cudaMemGetInfo` and a
one-line prefill estimate (`160 + chunk * 680 / 1024`, ~7x high at a
24576-token chunk), and only found out at the first prompt whether the chunk it
had committed to could still be served - the reported 524K failure: an int8 KV
of ~3.7 GiB, a cache that shrank around it, and a `--prefill 24576` whose
buffers no longer fit anywhere, so the engine died between the model load and
READY.
Enky's measurements on a 5080 (Windows/WDDM) sharpened this: an allocation that
does not fit does not necessarily fail - under default WDDM the driver pages it
into shared system memory (+9.25 GiB of Shared Usage, prefill collapsed to
~105 tok/s), so "cudaMalloc succeeded" proves nothing about the configuration
being safe.
## What this does
- **One startup VRAM plan** (`strata/prefill/vram_plan.{hpp,cpp}`): the prompt
path is priced with the exact `Prefill::bytes_needed()` BEFORE the expert
cache is committed; expert residency is the elastic consumer and gives way
first; an impossible configuration is refused at startup with its budget and
the knobs that make room, instead of dying at the first prompt.
- **The chunk/loan policy exists once** (`plan_lend_chunks`): the #583
bisection with its ring what-ifs, 0.1.39's list, a fixed chunk's halving -
shared by the startup plan, the WDDM post-touch revalidation and the runtime
loans.
- **The accepted plan is the runtime's contract** (`EffectivePrefillPlan` +
`check_runtime_prefill_use()`): generate and single-GPU serve consume the
accepted chunk/loan/ring directly and refuse - before any prompt allocation -
on a mode or size departure. The runtime may run less (a shorter prompt),
never more.
- **WDDM post-touch correction**: every opened cache is touched, read and
revalidated; a shrink re-prices the prompt path against the final physical
cache (a borrowed plan that lost its loan is repriced as owned); the retry
limit bounds shrinking, never validation. The loan's request-time
bookkeeping (`PfPart`) is built from the accepted plan, so lend/refill keep
the cache residency contract.
- **One borrowing capability** (`prefill_borrow_available`): the residency map
is built whenever the prompt path may borrow, token graph or not, and
startup and runtime read the same flag.
- **Enky's findings**: `--prefill auto` owned fallbacks are logged (planned vs
unexpected); the serve pre-check refuses when no owned chunk through 512
fits instead of handing the allocation to WDDM (with a Windows sysmem
warning); owned fallbacks reduce on the 256-token grid; budget hints name
`--kv-resident` as a lever where the model supports KV streaming.
## Measured (RTX 4070 Ti SUPER 16 GB, Linux)
| arm | before | after |
|---|---|---|
| 32K, `--prefill auto` (IQ3_XXS, int8 KV) | 8192 borrowed / 2489 slots / 1946 tok/s | same plan, 1906–1933 tok/s |
| 524K, int8 KV, `--prefill 24576` | died between model load and READY | 17664 owned, 2139 tok/s prefill, decode 22.8 tok/s |
| 524K, no KV residency, `--prefill 24576` (the WDDM-paged config) | unsafe 24576 handed to the driver | 2048 owned, selected at startup, 1104 tok/s |
| serve, two 8K requests back to back | - | each lend marks 2489 experts, each refill restores 2489, no stale residency |
## Tests
- `vram_plan_test` (new, no GPU): the planner's budget cases, the shared
chunk/loan policy, the runtime contract, Enky's fallback reductions (E1–E5)
and the runtime-mode cases (R19–R34) - all pass.
- `ctest`: same results as main (the 9 pre-existing environment failures are
unchanged); `prefill_fused_iq_test` passes.
- Existing behavior preserved: `STRATA_RING_BYTES`, `STRATA_PREFILL_RING`,
`STRATA_PREFILL_LEND_PCT`, the 128-slot loan floor, `--expert-cache-per-layer`,
`--vram-elastic`, the reserve adaptation on small cards (#496), and multi-GPU
stage planning (untouched by design - the planner does not model stage caches
yet).
## Not covered
- Windows/WDDM validation with "Prefer No Sysmem Fallback" (no Windows machine
available); the planner decides from arithmetic before any allocation, so
the fallback setting has nothing to change - but the issue's 16-GiB
Blackwell arm deserves a re-run on the original rig.
- Layer-split stage caches are still reserved with the split's own estimate
(sessions for later stages do not exist at planning time); their per-stage
serve logic is unchanged.Mehr auf der Site
Links zu Install, Modellen, Releases.