贡献 / #1704
#1704 prefill: opt-in active-layer cache, P4 108.1 -> 115.8 tok/s
open · @CC-David-CC · 0 评论 · 去 GitHub 看
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
说明
Fresh prompts repeatedly transfer the same nonresident expert weights. This adds an opt-in, VRAM-bounded cache for the active layer, keeping its weights across ordinary chunks in a bounded token window. Enable with `STRATA_PREFILL_LAYER_CACHE=1`; default remains off. One Tesla P4, Qwen3.8-Flash-Next-GSQ-RCO IQ3_XXS, automatic chunk sizing, 16K context, int8 KV, MTP spec 2, 64 generated tokens and no prefix reuse: | Prompt / sample | Off tok/s | On tok/s | Gain | | --- | ---: | ---: | ---: | | 4K, median of 2 per arm after MTP warmup | 108.133524 | 115.816006 | 7.105% | | 8K, initial single pair | 115.079293 | 122.610686 | 6.545% | 4K individual runs: off 107.881741 / 108.385308; on 115.447274 / 116.184738. Initial 4K pair: 108.037391 -> 114.704655 (+6.171%). Initial expert H2D bytes: 4K 105,829,222,400 -> 53,896,192,000; 8K 193,275,468,800 -> 98,652,032,000. Initial and steady-state warmups differ; samples are reported separately. Small samples on one workload. Uses existing per-chunk kernels, pinned-host residual windows and deterministic expert-ID admission. Remaining experts stream normally. Adds a coherent-window checkpoint guard and a reproducible benchmark harness. No new GPU kernels, server changes, persona profiles or prefix-pinning policies; global decode-cache placement remains unchanged. Validation: all compared 64-token outputs exactly match, including the unmodified main reference; sampled final residuals are bit-identical. All 8 checkpoint integration requests passed (restore, branching, protected restart and deleted-cache replay). Cancellation recovers on the same engine. Companion Python suite: 96 passed, 2 optional-dependency skips. Inspired by the transfer-reuse discussion in #1670; independently implemented using existing kernels. Developed and integration-tested alongside #1667, then extracted onto main with identical native source. Neither PR depends on the other. Measurements, configurations, raw timings, exact token IDs and limitations: [validation record](https://github.com/CC-David-CC/Strata-a5500/blob/perf/prefill-active-layer-cache/docs/measurements/p4-layer-cache/README.md). Draft: CUDA built and exercised on one P4. HIP/SYCL builds remain outstanding before review, per AGENTS.md. Multi-GPU, images and unsupported layouts decline the experiment; AMD, RTX 3070 and larger contexts were not validated.
本站相关内容
相关页面的快捷入口。