Pull requests / #1637
#1637 batch: the adaptive expert tier keeps adapting in batch windows (`--adapt-async 1` beside slots)
open · @noon-at-cgn · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
本文
## Title Issue: Related: #876 (the asynchronous tier), #1122 / #905 (pipelined windows with the tier), #1413. Depends on #1636 (`--batch-mtp` on a layer split, which adds the `--batch-groups` resolution this uses). ## Summary **Why it matters:** on upstream main the VRAM expert tier does not adapt in batch windows at all: `batch_step` never calls it, `--adapt-async 1` is switched off with "not with --batch slots", and pipelined groups never adapt. With `parallel 2` the cache then follows the routing only while a request is alone. This change counts the routing of batch windows and runs the tier every `--adapt-every` windows, blocking by default and between the windows with `--adapt-async 1`, on one GPU and on a layer split (every stage's cache adapts). On 0.1.41 it needs one more thing: a layer split with `--batch` 2 or more pipelines the slots by default and the pipelined path never adapts, so with `--adapt-async 1` and no `--batch-groups`, `auto` resolves to one group and the log says why; an explicit `--batch-groups N` or `auto` is honoured and turns `--adapt-async` off. | Measured: 2 concurrent greedy decodes of 300 tokens, layer split `23`, `parallel 2`, `--batch-mtp`, 2x RTX 3080 20 GB (220 W cap), UD-Q4_K_XL, engine 0.1.41 = our stack on fb58e0d (strata-w10) | aggregate tok/s | per window | routed entries from VRAM | |---|---|---|---| | `--adapt-async 1 --adapt-every 2` (3 restarts, 60 measurements) | **89.2** | 34.6 ms, adapt wait 0.2 ms | 92.1% | | blocking tier, `--adapt-every 2` (1 restart, 20) | 74.7 | 41.4 ms, adapt wait 8.4 ms | 92.3% | | no tier, the cache frozen from a profile the tier had learned on the same prompts (1, 20) | 69.5 | 43.6 ms | 77.2% | | no tier, `--adapt-every 0` (2 restarts, 40) | 45.2 | 66-68 ms | 52.7% | Read it as the tier as a whole, not as "batch windows only": the switch also changes the solo path (solo decode: 81.7, 66.3, 69.0, 43.1 tok/s in the same order), and no switch turns the tier off in batch windows alone. The prompts are the benchmark's own (eight topics that repeat), which the tier learns quickly, so on a varied workload the gain may be smaller (not measured). The async tier waits 0.2 ms per window, the blocking tier 8.4 ms. Upstream main runs no tier in batch windows and no upstream binary was measured. ## What changed - `src/program/generate.cpp`: in `batch_step`, windows count the routed experts (`drive.d.usage`); every `--adapt-every` windows the blocking tier runs on a thread beside the commit and the drafts, and its swaps land (`apply_pending`) before the window ends; with `--adapt-async 1` the round in flight moves on one step per window (table changes and uploads on the main thread, copies on the helper thread and the cards' refill streams). Not between a prompt's chunks (a loan of cache slots), and a request that arrives lands the round in flight first. The `--adapt-async` gate says "not with --batch-groups (the pipelined slot groups)" instead of "not with --batch slots". The asynchronous round ends at once when its swap list turns out empty. `--adapt-min-gain F` (default 1.5, the old constant) sets the bar a swap must clear, for both tiers. The `strata batch:` line shows each stage's GPU-reach wait and pool time, the routed entries served by VRAM, PCIe and CPU, and the rounds and swaps. - `include/strata/program/batch_groups.hpp`, `src/program/batch_groups_test.cpp`: `--adapt-async 1` joins `--batch-mtp` as a reason the default `auto` gives way to one group (log: `--batch-groups auto: 1 group of 2 slots - --adapt-async 1 does not run in pipelined groups; --batch-groups 2 would pipeline them and turn it off`, read in the start log of arm A1, no `--batch-mtp`). With `--batch-groups 2` given, the engine printed `--adapt-async 1 is off (not with --batch-groups (the pipelined slot groups))`. - `docs/BATCHING.md` (a section and the table above), `docs/DETAILS.md`. - Commits (all by noon-at-cgn): the tier in batch windows, the asynchronous tick with `--adapt-min-gain`, the groups resolution for `--adapt-async 1`, the docs with the measurements. ## Extra Notes **Behaviour change to know about:** the blocking tier now runs in batch windows by default (every 4 windows, `--adapt-every`), where upstream main runs none; `--adapt-every 0` (or `1000000`, as in the exactness settings of docs/BATCHING.md) keeps the tier fixed. Swaps move experts between the GPU and the CPU, which round differently, so greedy outputs of a batch can differ from run to run when the tier is on; with it off the exactness statement of docs/BATCHING.md is unchanged. The measured gain above is for a box where most experts are resident in RAM and the VRAM tier is a cache of them (`--resident-experts`); `--adapt-async 1` needs that mode. **Order:** #1636 first. This branch applies on top of it without conflicts and was built and tested on top of it (not on main alone). It does not conflict with open #1598, #1599, #1601, #1614 or #1517 (checked with `git merge-tree`). It touches the same adaptive-tier code as #1517 (evictions reach the residency table before the next window) without a textual conflict; the two have not been run together. **Not tested:** HIP, SYCL, three or more stages, more than two slots, images with batch slots, sampled decoding, GPU parity of the swap path in batch windows (no exactness claim: the tier is not bit-exact run to run), `--adapt-async 1` with `--pipeline-windows 2` beside slots, `--peer-device`, the helper caches, any upstream-main binary. The numbers come from our build with the rest of our stack (shared KV pool, #1190, aux-cpus, memory guard); the tier arms were run in alternation with the batch-MTP arms between 15:53 and 20:34 on one day, with 1 to 3 restarts per arm. Tests: `batch_groups_test` (the 0.1.41 default, `--batch-mtp` and `--adapt-async 1` giving way, explicit values honoured; 16 checks) and `batch_rows_test` pass; `ctest` on the whole tree with no GPU: the same 56 failures as #1636 (missing GPU). On GPU 0 with about 1 GiB free: `verify_parity`, `verify_batch_parity`, `rope_parity`, `qsa_parity`, `kv_stream_parity`, `kv_hybrid_parity` pass. `python -m unittest discover -s serve`: 593 tests OK (11 skipped). The adaptive code path itself ran only in the production engine on strata-w10 (the arms above); there is no CPU test of it. Context: this change is part of the setup measured in the community benchmark #1640.
関連リンク
インストール・モデル・リリースへの站内リンク。