Issues / #765
#765 Make VRAM budgeting aware of KV, prefill staging and expert residency
open · @j-luwierski · 4 comments · View on GitHub
Description
Problem
At very large context sizes, the current VRAM budgeting order can leave too little device memory for prefill staging buffers even though the model itself can otherwise fit and run.
This was observed in the 524K context testing reported in #760.
With:
--rope-scaling yarn
--rope-scale 2.0
--max-context 524288
and int8 KV, the KV pool consumes roughly 3.7 GB of VRAM.
The KV allocation happens before the automatic expert-cache sizing settles on the remaining device budget. As a result, the main GPU can end up with substantially reduced expert residency. That part is relatively benign by itself, but the more important consequence is that the remaining free VRAM is no longer sufficient for large prefill staging buffers.
For example:
--prefill 24576
requires roughly 2.4 GB of temporary/device-side prefill buffers.
At 524K, the engine can then fail after model loading but before reaching "READY", because those buffers can no longer be allocated.
Reducing the chunk size to:
--prefill 4096
allows the engine to boot, but comes with a significant performance cost. In the reported A/B test, prefill throughput dropped from approximately:
2159 t/s -> 1778 t/s
at 256K, or about 18%.
The obvious workaround, increasing:
--vram-reserve-mib
does not solve the underlying problem either. Reserving more VRAM for later allocations can instead starve the weight arena, which needs roughly 1.4 GB, causing startup to fail earlier.
So there is currently no single VRAM control that can reliably express:
«preserve enough memory for the weight arena and prefill staging first, then use the remainder for expert residency.»
Root cause
The problem appears to be primarily one of allocation order and independent VRAM budgeting decisions.
Several large consumers compete for the same device memory:
model / weight arena
KV cache
prefill staging buffers
expert cache / expert residency
other runtime buffers
but they are not budgeted together before allocations are committed.
At large contexts, KV usage grows enough that a configuration which looks acceptable during one stage of initialization can become impossible later when prefill staging buffers are allocated.
The auto expert-cache calculation also cannot make the best decision if it does not know how much memory must still be reserved for prefill staging and other mandatory runtime allocations.
Proposed solution
Introduce a unified VRAM budgeting step before the large optional allocations are committed.
The allocator should calculate or estimate the mandatory VRAM requirements first:
fixed runtime overhead
+ weight arena
+ KV cache
+ prefill staging requirement
+ safety margin
and only then assign the remaining VRAM to optional / elastic consumers such as expert residency.
Conceptually:
available_vram
- fixed_runtime
- weight_arena
- kv_cache
- prefill_staging
- safety_margin
= expert_cache_budget
1. Reserve prefill staging before sizing the expert cache
When automatic expert-cache sizing runs, it should know the expected staging requirement for the selected prefill chunk size.
Expert residency should be reduced before allowing the staging allocation to fail.
Since reduced expert residency mainly affects performance while failure to allocate staging prevents the server from starting at all, this seems like the preferable tradeoff.
2. Make "--prefill auto" memory-aware
If the requested/default prefill chunk does not fit even after reducing elastic expert residency, automatically choose the largest chunk that satisfies the VRAM budget.
For example, the selection could try:
24576
16384
12288
8192
4096
...
or use a direct calculation if staging-memory requirements are predictable enough.
The selected value should be logged explicitly, for example:
prefill: requested 24576, clamped to 8192 due to VRAM budget
This would be much better than failing between model load and "READY".
3. Treat expert residency as elastic
Auto expert-cache sizing should operate on the VRAM that remains after mandatory allocations are accounted for.
In other words, expert residency should not consume memory that will later be required by:
- weight arena,
- KV cache,
- prefill staging,
- mandatory runtime scratch buffers.
This may reduce the number of resident experts at very large contexts, but that degradation is graceful and measurable rather than fatal.
4. Improve startup diagnostics
If no valid configuration fits, report the VRAM budget instead of a generic allocation failure.
For example:
VRAM budget:
available: 15.4 GiB
weights/runtime: X.X GiB
weight arena: 1.4 GiB
KV cache: 3.7 GiB
prefill staging: 2.4 GiB
expert cache: X.X GiB
safety margin: X.X GiB
Unable to satisfy requested --prefill 24576.
Try a smaller prefill chunk, lower KV precision, or a smaller context.
This would make the problem much easier to diagnose.
5. KV compression remains an additional lever
For extreme context sizes, reducing KV bytes is still useful and may be necessary on 16 GB cards.
For example, a configuration such as "k8v4" may free enough memory to retain a larger prefill chunk.
However, KV compression should be considered an optimization / user choice rather than the only workaround for an allocation-order problem.
Expected behavior
For configurations that can fit by reducing expert residency or prefill chunk size:
- the server should reach "READY",
- mandatory runtime allocations should never be starved by automatic expert-cache sizing,
- "--prefill auto" should choose the largest safe chunk,
- the selected budget and any automatic clamping should be visible in the logs.
A hard startup failure should only occur when no valid combination fits within the available VRAM.
Suggested tests
Regression test: budget ordering
Simulate or constrain available VRAM so that:
KV + weights + large expert cache + requested staging > available VRAM
but:
KV + weights + smaller expert cache + requested staging <= available VRAM
Expected result:
expert residency is reduced
server reaches READY
requested prefill chunk is retained
Prefill fallback test
Create a case where the requested prefill chunk does not fit even with minimum acceptable expert residency, but a smaller chunk does.
Expected result:
requested prefill is automatically reduced
server reaches READY
warning/log explains the clamp
Hard-failure test
Create a case where even the minimum supported prefill configuration cannot fit.
Expected result:
startup fails cleanly
VRAM budget is printed
error identifies the limiting allocations
Large-context hardware validation
Re-run at 384K / 512K / 524K on 16 GB GPUs and record:
- selected prefill chunk,
- KV allocation,
- expert residency,
- weight-arena allocation,
- free VRAM before "READY",
- prefill throughput,
- decode throughput.
The original 524K report from #760 would be a useful baseline for this validation.
Related
Discovered during the 524K consumer Blackwell testing in #760.Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.