Pull requests / #294
#294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windows
closed · @giostrives · 0 comentários · No GitHub
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
This PR is scavenged from https://github.com/Niko1221/Strata/pull/223, limiting the changes to the memory pinning and multi-gpu planner. - Moves the `--layer-split auto` cost model out of `generate.cpp` into a header-only, unit-tested `Planner` - Adds a hand-off cost to the model and stops capping the expert-arena pin on Linux. ## Changes - **Planner extracted** (`include/strata/program/layer_split.hpp`). The placement arithmetic is now in `layer_split::Planner`, `Card`, `Costs` and `Placement`, apart from the devices. `generate.cpp` only gathers per-card VRAM and SM/clock data and calls `planner.best(cards)`. - **Hand-off term.** The predicted decode window now includes one hand-off per later card (`STRATA_SPLIT_HANDOFF_MS`, default 1.0 ms; about 1.1 ms measured on 2 x RTX 2080 Ti over PCIe gen3). The startup log's `predicted ... ms per decode window` includes it. - **New tunables.** `STRATA_SPLIT_HANDOFF_MS` and `STRATA_SPLIT_MASS_EXP` join `STRATA_SPLIT_MISS_MS`. Defaults are unchanged from the previous hard-coded values (190 ms, exponent 1.2). - **Arena pinning.** `arena_pin_cap(several_contexts, wddm)` caps the pin at 8 GiB only under WDDM with several contexts. Linux now pins the whole arena, as with one card. - **Tests.** New `layer_split_test` (registered in CMake) covers the pin cap, per-layer time, the fill-stops-at-first-misfit rule, hand-off terms, placement search for 2, 3 and 4 cards, and the carve pricing. - **Docs.** `docs/MULTI_GPU.md` describes the new model, the per-layer sessions and borrowed prompt buffers, and adds a 2 x RTX 2080 Ti measurement. ## Measurements Qwen3.8-Flash-Next IQ2_XS (33 GiB of experts, more than both cards hold), 2 x RTX 2080 Ti (11 GB, PCIe gen3 x16, i7-7700K), Linux, 32K context, `--spec 4`. Auto chose K=25. | | Prompt 1.9K / 8K / 16K / 29K tok/s | Decode tok/s | Decode cache hits | |---|---|---|---| | one 2080 Ti | 466 / 695 / 840 / 836 | 36.7 | 66% | | 2 x 2080 Ti, auto (K=25) | 485-490 / 774-781 / 1,116-1,131 / 1,278-1,302 | 48.4-50.3 | 80-81% | | same, arena pinned only to 8 GiB | 442 / 696 / 921 / 999 | 44.7 | - | Decode is +32-37%. Prompts gain with length: +4-5% at 1.9K, +53-56% at 29K.
No site
Links install, modelos, releases.