Issues / #658

#658 Proposal: single-GPU prefill buffer planning and RAM budgeting for conversation caching (RTX 5090 measurements)

open · @spideytznn · 1 commentaires · Sur GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentationWindows

Description

## 中文摘要

我们在基于 v0.1.35 的 Windows / RTX 5090 本地分支中保留了两项小规模内存规划改造,想分享实现和测量结果,供上游评估:

1. 单卡专家缓存填满后,仅用剩余显存选择独立预填充缓冲;保留显存余量,批量不小于借用路径,不主动缩小专家缓存。
2. 启用 RAM 常驻专家及上游多会话缓存时,将会话预算、最低空闲 RAM 和 256 MiB 余量纳入专家 headroom。

一次同二进制顺序对照中,17944-token 完整预填充从 2311.6 到 2901.7 tok/s(吞吐约 +25.5%,耗时约 -20.3%)。默认自动专家缓存下,两版约 2925 / 2919 tok/s,基本一致。单轮测量有文件缓存和内存状态影响,不能作为普遍提速结论。会话缓存、检查点和恢复实现全部来自上游。

详细实现、条件、验证和限制如下。测量基于 v0.1.35,尚未在 v0.1.38 上重新验证。

## Proposal and implementation

I maintain a small local fork for a Windows RTX 5090 system. Two memory-planning changes may be useful upstream: choosing independent prefill buffers when spare VRAM can pay for them without reducing the expert cache, and reserving conversation-cache RAM when planning resident experts.

The implementation is based on upstream **v0.1.35**, commit `d9ab8435f654c368c586340d490915f6addf56a3`:

- [Implementation commit and complete diff](https://github.com/spideytznn/Strata/commit/b02e3b7ba9dcc0efbd697d6920615002e918fd85)
- [Pure planning helpers](https://github.com/spideytznn/Strata/blob/b02e3b7ba9dcc0efbd697d6920615002e918fd85/include/strata/program/local_memory_plan.hpp)
- [Integration in generate.cpp](https://github.com/spideytznn/Strata/blob/b02e3b7ba9dcc0efbd697d6920615002e918fd85/src/program/generate.cpp)
- [11 boundary checks](https://github.com/spideytznn/Strata/blob/b02e3b7ba9dcc0efbd697d6920615002e918fd85/src/program/local_memory_plan_test.cpp)
- [Bilingual implementation and validation notes](https://github.com/spideytznn/Strata/blob/0c824e173f0873993b5fd83db570c03537127ac0/docs/LOCAL_VARIANT.md)

I checked the v0.1.38 source: these two local planning policies are not present there. **All measurements below are on v0.1.35; the patch has not been rebased or benchmarked on v0.1.38 yet.**

### 1. Prefer independent prefill buffers when spare VRAM permits

Borrowing expert-cache slots can require restoring their experts, and in RAM-constrained situations some lent experts may need file reads. The local planner tries to avoid lending when an independent buffer fits:

- Run after the existing GPU expert cache is filled; query `cudaMemGetInfo`.
- Price candidates using `Prefill::bytes_needed(g, ss, chunk)` for chunks `32768, 16384, 8192, 6144, 4096, 3072, 2048, 1024, 512, 256`.
- Require `buffer_bytes <= free_VRAM - configured_VRAM_reserve` and respect the requested/context chunk cap.
- Choose the largest fitting candidate **at least as large as the chunk selected by upstream's borrowing planner**.
- Keep the existing expert-cache size. If no candidate fits, retain the upstream borrowing path.
- When selected, skip lending those cache slots for prefill and skip reserving RAM copies specifically for the would-be lent slots.
- Apply only to single-GPU local-cache execution with an expert profile. Layer splits and remote-cache paths retain upstream behavior. Explicit `--no-prefill-borrow` retains its upstream meaning.

`STRATA_PREFILL_OWN_AUTO=0` disables only this local automatic choice for same-binary A/B comparisons. The planner preserves the reserve at its decision point; that is not a guarantee against later allocations or other processes consuming VRAM.

### 2. Account for upstream conversation caching in resident-expert RAM planning

Resident experts and parked conversations compete for the same physical RAM. Before resident-expert allocation, the local helper computes:

```text
expert_headroom = max(existing_headroom,
                      conversation_cache_budget
                    + conversation_cache_min_free
                    + 256 MiB)
```

This applies only when resident experts and serving conversation caching are enabled, including nonzero prompt-cache and conversation-slot limits. Larger existing headroom is respected; disabled/zero-budget conversation caching leaves upstream headroom unchanged. The helper saturates on integer overflow.

For `--conversation-cache-mib 2048 --conversation-cache-slots 2 --conversation-cache-min-free-mib 2560`, the required headroom is **4864 MiB (4.75 GiB)**. This is a planning budget; upstream's physical-RAM admission check remains necessary, and other applications can still consume memory later.

Snapshot accounting, prefix checkpoints, restore, isolation, growth, and eviction are **upstream implementations**, not work introduced by this patch. There is no SSD session cache and no new concurrent-inference implementation.

## Measurements

Environment: Windows, RTX 5090 32 GiB VRAM, 48 GiB system RAM, CUDA 13 / MSVC Release (`sm_120`), Qwen3.8-Flash-Next IQ3_S native GGUF, FP8 PLE, INT8 KV, MTP, 262144-token context limit. These comparisons call the engine directly without the separate vision helper.

### Same-binary buffer-policy comparison

Both arms used the local binary, `--prefill auto`, `--vram-reserve-mib 1536`, the same profile and conversation-cache settings above, and `STRATA_RESIDENT_HEADROOM_GIB=4.75`. Only `STRATA_PREFILL_OWN_AUTO=0` versus `1` changed. `--expert-cache 5400` was used to leave spare VRAM; the engine reported **7056 slots / 13709 MiB (13.39 GiB actual expert-cache allocation)** in both arms. The actual allocation is what matters here; 5400 is the requested CLI value, not the measured allocation.

| Warm-file full prefill | Borrowing | Independent |
| --- | ---: | ---: |
| Input tokens / reused tokens | 17944 / 0 | 17944 / 0 |
| Batch size | 8192 | 8192 |
| Prompt time | 7762.6 ms | 6183.9 ms |
| Prompt throughput | 2311.6 tok/s | 2901.7 tok/s |
| Expert slots | 7056 | 7056 |

This is **+25.5% throughput / -20.3% prompt time for this pair**. The independent buffers were 3805 MiB. The borrowing path lent 1972 slots and logged 1799 additional file-blob reads during each long prefill; the independent path logged zero file-blob reads during those requests. Those observations are consistent with the intended mechanism, but do not isolate it from all other effects.

This was **one sequential pair, borrowing first**, with warm-file/order and physical-memory confounding; resident RAM allocation also differed (33.78 versus 33.45 GiB). It is preliminary evidence for the option, not a general speedup claim or a pure GPU-kernel microbenchmark. Arithmetic returned `391`, and both long prompts returned `OK`.

To reconstruct the synthetic workload, render a single user message with thinking disabled and temperature 0. The first long prompt is `Read the following records. At the end answer only OK.\n` followed by newline-separated records generated with:

```python
document = '\n'.join(
    f'Record {i:04d}: The warehouse has blue boxes, green triangles and red circles. '
    'Check the count and retain the label.'
    for i in range(640)
)
```

The measured warm prompt uses `Separate document. Read it carefully and answer only OK.\n` plus the same document, forcing a full prefill rather than prefix reuse. Each engine arm first answers `Answer only the number: 17 times 23.`, then the first long prompt, then the measured distinct long prompt; the long requests allow up to 16 output tokens. The engine reported 17944 input tokens for both long prompts with the tested tokenizer/template.

### Normal auto-sized cache: no established speed improvement

The production configuration keeps `--expert-cache auto`; it was not reduced to obtain the result above. A separate official/local v0.1.35 comparison used equal 2 GiB conversation budgets, explicit 4.75 GiB headroom, and equal 10222 expert slots. Both chose borrowing. Warm full prefill was **2924.7 versus 2919.0 tok/s**: effectively unchanged in this short comparison.

One 512-token decode pair was 121.2 tok/s official versus 112.6 local. Three additional runs in reverse engine order measured 107.3/146.0/136.7 local versus 95.7/94.2/128.2 official. This spread does not establish a stable decode improvement or regression. No speed benefit is separately attributed to the RAM-budget helper.

## Validation and upstream consideration

- 11 local planner boundary checks and 30 upstream conversation-cache Python tests passed.
- CUDA 13 / MSVC Release build for `sm_120` passed.
- Upstream real-model A→B→A caching parity checks passed on IQ3_S / INT8 KV / FP8 PLE: output tokens and restored main-model state hashes matched with caching off/on. One 984-token checkpoint restore took about 26 ms; that restore functionality belongs to upstream.
- A separate complete-service smoke check passed with GPU vision and HTTP A→B→A reuse. Four upstream queue/cancellation/handover tests passed with mock engines; these are not a real-model concurrency stress test.

Would these policies be useful upstream, perhaps initially as an opt-in? The main follow-up would be rebasing onto v0.1.38 and running alternating repeated comparisons across cache sizes, RAM pressure and prompt lengths. I would especially appreciate feedback on the VRAM decision point and the conversation-headroom formula.

Sur le site

Liens install, modèles, releases.