Pull requests / #1458
#1458 helper cache: configurable VRAM allowance with allocation-floor checks
open · @saikiran-rs · 0 comentarios · En GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Descripción
Helper expert caches use a fixed 512 MiB allowance during admission. Add `STRATA_REMOTE_RESERVE_MIB` so a measured setup can explicitly choose a different allowance, while retaining 512 MiB when unset. ## What changed - Apply the allowance to automatic sizing and explicit helper slot counts. - Require space for the existing device work buffers plus 16 MiB of headroom before filling the cache. - Reject malformed, out-of-range and undersized values. Admission arithmetic uses subtraction to avoid overflow. - Add a CPU-only policy test covering the unchanged default, parsing, the buffer floor and capacity boundaries; document the setting in SECOND_GPU.md. This changes helper cache admission only. The historical 128 MiB setting was used on one Linux RX 7900 XTX + RX 6800 XT machine; the default remains 512 MiB. Its combined tuning results are documented in the [companion benchmark report](https://github.com/saikiran-rs/Strata/tree/community/amd-helper-1m-report-20261007/bench/results/2026-10-07-community-xtx-6800xt-1m-helper), with no isolated allowance speedup claimed. ## Validation Based on upstream `d5ea7133741e67743c0e886bb426c0ce8d69cf6c` (0.1.40.3): - CPU-only CMake build and `ctest -R '^remote_reserve_test$'`: passed. - Changed `remote_experts.cpp` compiled against the HIP SDK in a fresh isolated build configuration: passed. - `git diff --check`: passed. No new GPU/model inference run was performed on this isolated PR. Historical full-context/quality results refer to the earlier combined research build. CUDA and Windows runtime testing remain unmeasured. Related #1184 handles primary-device selection and resident-RAM helper support; #848 handles a layer split's reserve. Their changes are not duplicated here.
En el sitio
Enlaces a install, modelos, releases.