Pull requests / #1704

#1704 prefill: opt-in active-layer cache, P4 108.1 -> 115.8 tok/s

open · @CC-David-CC · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Beschreibung

Fresh prompts repeatedly transfer the same nonresident expert weights. This adds an opt-in, VRAM-bounded cache for the active layer, keeping its weights across ordinary chunks in a bounded token window. Enable with `STRATA_PREFILL_LAYER_CACHE=1`; default remains off.

One Tesla P4, Qwen3.8-Flash-Next-GSQ-RCO IQ3_XXS, automatic chunk sizing, 16K context, int8 KV, MTP spec 2, 64 generated tokens and no prefix reuse:

| Prompt / sample | Off tok/s | On tok/s | Gain |
| --- | ---: | ---: | ---: |
| 4K, median of 2 per arm after MTP warmup | 108.133524 | 115.816006 | 7.105% |
| 8K, initial single pair | 115.079293 | 122.610686 | 6.545% |

4K individual runs: off 107.881741 / 108.385308; on 115.447274 / 116.184738. Initial 4K pair: 108.037391 -> 114.704655 (+6.171%). Initial expert H2D bytes: 4K 105,829,222,400 -> 53,896,192,000; 8K 193,275,468,800 -> 98,652,032,000. Initial and steady-state warmups differ; samples are reported separately. Small samples on one workload.

Uses existing per-chunk kernels, pinned-host residual windows and deterministic expert-ID admission. Remaining experts stream normally. Adds a coherent-window checkpoint guard and a reproducible benchmark harness. No new GPU kernels, server changes, persona profiles or prefix-pinning policies; global decode-cache placement remains unchanged.

Validation: all compared 64-token outputs exactly match, including the unmodified main reference; sampled final residuals are bit-identical. All 8 checkpoint integration requests passed (restore, branching, protected restart and deleted-cache replay). Cancellation recovers on the same engine. Companion Python suite: 96 passed, 2 optional-dependency skips.

Inspired by the transfer-reuse discussion in #1670; independently implemented using existing kernels. Developed and integration-tested alongside #1667, then extracted onto main with identical native source. Neither PR depends on the other. Measurements, configurations, raw timings, exact token IDs and limitations: [validation record](https://github.com/CC-David-CC/Strata-a5500/blob/perf/prefill-active-layer-cache/docs/measurements/p4-layer-cache/README.md).

Draft: CUDA built and exercised on one P4. HIP/SYCL builds remain outstanding before review, per AGENTS.md. Multi-GPU, images and unsupported layouts decline the experiment; AMD, RTX 3070 and larger contexts were not validated.

Mehr auf der Site

Links zu Install, Modellen, Releases.