Pull requests / #559
#559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
closed · @blange48 · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows
Beschreibung
## Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights Strata serves one request at a time today (#465, #249): with several clients everything queues. This PR lets one engine process decode **several conversations together**, opt-in, without changing anything when the new options are absent. Full description in the new `docs/BATCHING.md`. ### What it adds **1. Batch windows in the verifier (`--batch N`, 2..8)** - `include/strata/core/verify.hpp`, `src/core/verify.cpp` - A batch window holds one token of each of S independent sequences. Row `s` is slot `s`, at its own position, reading and writing its own state, which lives in its own `SessionState` (carved with `session_init` like the stage's own session: same layer range, same `max_cells`). - Only the per-sequence parts of `record_window` loop over the slots: the GDN conv history and recurrence (`gdn_conv_l2_multi` / `gdn_step_norm_multi` on each slot's state, one row each), the QSA K/V append, indexer append, block selection and attention on each slot's `QsaState`, and the PLE history. Everything that is per row already - hyper-connections, the dense projections, the router, the shared expert, the routed experts on the GPU and the CPU pool, the head - runs once over the S rows, so the weights are read once per window for all of them. - A batch window keeps every row, so its commit (GDN state, indexer tail, PLE history, per slot) is a captured graph that needs no host decision. - New API: `init_slots`, `run_slots` / `commit_slots` (synchronous), and `batch_launch` / `batch_poll` (asynchronous, for the pipeline below). Graphs are captured per (first slot, count). - Sampling per slot (temperature / top_p / top_k / min_p / seed), drawn row by row on the last stage with the solo window's Philox(seed, position). **2. A pipelined layer split (`--batch-groups G`)** - `src/program/generate.cpp` - Today the stages of a layer split run one after the other: one card works, the others wait. With G groups of slots, a batch window over one group is launched on one stage with its commit right behind it on the stage's stream, and one host thread serves the rings of every stage that has a window in flight (`batch_poll` does not block). Stage k then runs group g while stage k+1 runs group g-1. - Each group uses its own rows of the existing hand-off buffers, so no new buffers are needed. **3. Per-stage dense weights (`--trim-stage-weights`)** - `src/program/generate.cpp`, `src/core/native_dense.cpp` - With an explicit `--layer-split`, every GPU loaded the whole model's dense weights although it runs a quarter of the layers. With the option, each GPU skips the `blk.<l>.` tensors outside its range, both in the pack (`WeightTable::load`'s skip list) and in the native projections (`NativeDense::load` takes a layer range). The VRAM goes back to the expert cache. Useful without `--batch` too. **4. The engine protocol and the server** - `src/program/generate.cpp`, `serve/server.py` - `BGEN <slot> <max_new> [keys] <ids>` (and `BGENI`) reads the prompt through the usual path as a `GEN 1` (prompt cache, checkpoints, images), then copies the state into the slot with the existing `conversation_checkpoint_save/restore` and `conversation_kv_save/restore` on every stage, and answers `BADM <slot> 1|0`. Slots stream `BT <slot> <id>` and end with `BDONE <slot> ...`; `BSTOP <slot>` ends one. Between batch windows the engine reads its input without blocking. - The server (`StrataEngine.generate_batched`) runs concurrent requests in the slots when the engine has `--batch`: a request alone takes the solo path (MTP drafts, the fastest single stream); when another one arrives, it is STOPped and continues in a slot with its prompt plus what it generated (the prompt cache holds exactly that), and the newcomer is admitted next to it. The request FIFO is bypassed in this mode. **5. A client that stops reading early** (second commit) - `serve/server.py` - When a consumer leaves mid-answer (a disconnect, a tool call or a stop string ending the response), the solo request is STOPped and read to its DONE, an admission to its BADM, and a slot is BSTOPped and freed only at its BDONE. Without it the engine went on writing the abandoned answer and the next request read it as its own (seen with an agent client: the previous answer came back). `tools/early_close_test.py` reproduces it; `STRATA_BATCH_TRACE=1` logs the control lines. **6. Conversation parking with `--layer-split`** (third commit, by @crazyaimachine, from their comment below) - `include/strata/core/conversation_cache.hpp`, `src/program/generate.cpp` - A parked conversation holds one image per later stage; split checkpoints are cut into stage parts while parking and put back together on restore; every stage image is validated before anything is overwritten; the guard that rejected parking with a split is removed. - Verified here on a 4-GPU split (they tested 2): a 3,786- and a 9,276-token conversation restored in 42 / 51 ms, follow-ups token-identical to the conversation that never left (without parking: 3.4 / 6.8 s to re-read them). Through the HTTP server, two long conversations alternating: follow-ups at 0.53 / 0.46 s to the first token instead of 3.9 / 5.7 s. **8. The parking follow-ups** (fifth commit), the open points from that comment: - the draft layer's K/V is saved once, with the first stage's image: `conversation_snapshot_*` take a `const QsaState* draft`, `nullptr` for an image without it (the later stages); - every stage's restored K/V is retained (`ConversationKvReuse::stages`), so a conversation parked again copies only what changed on every stage - 141 MB reused on the second park of a 9K-token conversation in `parking_test.py`; - checkpoints are moved into their stage parts and back (`conversation_checkpoints_split` / `_merge`), not copied; - unit tests: `conversation_cache_test` (split/merge round trip without copies, refused misalignments, per-stage reuse) and `conversation_snapshot_test` (draft-less images: estimate, validation both ways, exactness, incremental reuse) - 4,191 and 2,155 checks passing. **9. The draft ring with the last stage** (sixth commit): the draft layer's K/V lives on the last stage's GPU (the drafter is loaded there), so it is now saved, restored and verified with that stage's image under that device, not with stage 0's on GPU 0 (where the ring-restore kernel ran against another GPU's memory - #653 places it the same way). Also: parking with a split on one GPU (`--split-device 0`) is refused at start, the whole-session validate refuses a split image, and `--trim-stage-weights` logs each GPU's layers. Re-verified on 4 GPUs with `STRATA_SNAPSHOT_VERIFY=1`: follow-up identical, draft ring read back, 141 MB of K/V reused on the second park. **10. MTP drafts in batch windows and active-slot group windows** (by @crazyaimachine, 9 commits from their comment below, applied unchanged): - a pipelined group window holds its slots only up to the last active one, instead of every slot of the group (idle slots ran pad rows through every layer); - `--batch-spec N`: MTP drafts in batch windows (rows per slot, a drafter per slot sharing the weights, per-slot `n_keep` commits, rows chosen per group window, drafts while at most `--batch-spec-max-active` slots are active), and the n-gram rows read token by token (a drift found on the way). - Re-verified on 4 GPUs: 8 slots identical to solo, 3 slots with `--batch-spec 2` identical, parking identical (141 MB of K/V reused on the second park). Through the HTTP server (T=0.7, 300 tokens): | Concurrent | before | active-slot windows | + `--batch-spec 2` (reserve 2000 MiB) | | ---: | ---: | ---: | ---: | | 1 | 106 | 106 | 97 | | 2 | 91 | **145** | 143 | | 3 | 149 | **201** | 194 | | 4 | 195 | **244** | 237 | | 8 | 354 | 359 | 346 | `--batch-spec` needs room for the slot drafters (at 262K context, 8 drafters did not fit in 700 MiB of reserve on 16 GB cards); the larger reserve costs expert residency, so on this machine it does not pay. **11. `tools/autoconfig.py`**: the settings above from the machine (GPUs, VRAM, PCIe link, RAM, layer count), with `--calibrate` to measure the candidates; on the 4-GPU machine the rules give back the hand-tuned config and name the card on an x8 link. Documented in `docs/BATCHING.md`. **12. `/metrics` for Prometheus, with vLLM's names** (three commits) - `serve/prometheus.py`, `serve/server.py`: - a scrape (`Accept: text/plain` / openmetrics, or `?format=prometheus`) gets the text format with vLLM's metric names (running / waiting, context usage, token and request counters, prefix-cache queries and hits, MTP drafts, TTFT / inter-token / end-to-end histograms), so dashboards and alerts written for vLLM read Strata; anything else keeps the JSON; - the JSON's own facts come in the same scrape under `strata:`, named after their JSON key (`strata:live_tok_s`, `strata:gpu_util` per card, ...), so the Monitor tab and Grafana read the same values; - with `--batch`, every concurrent request is now recorded (history, totals, latencies): they shared one status and only the last to finish was counted. `serve/test_prometheus.py` (6 tests, one with three requests at once). Documented in `docs/DETAILS.md`. **7. Tests** (fourth commit): `tools/parking_test.py`, `tools/batch_test.py --keys`, and a Testing section in `docs/BATCHING.md` listing the three scripts and how to run them exactly. Rebased on main (engine 0.1.38): the 8-conversation exactness test still passes there. ### Correctness A batch row's arithmetic is the single-token window's, so greedy outputs are the solo outputs: `tools/batch_test.py` decodes the same prompts alone and in batch and compares them token by token. **8 concurrent conversations x 150 tokens: identical to their solo runs**, with the synchronous batch and with the pipeline (2 and 4 groups), with `STRATA_IQ_MT_MIN=1` and `--pcie-frac 0`. With a PCIe share, which missed experts go to the GPU depends on the window's misses, so the rounding can differ and outputs drift apart after some tokens (as two solo runs with different windows can) - documented. ### Measured 4-GPU layer split (4 x 16 GB, PCIe Gen3), IQ3_S, `--batch 8 --batch-groups 4 --trim-stage-weights`, HTTP server, 400 tokens per answer, temperature 0.7: | Concurrent requests | Per request | Total | | ---: | ---: | ---: | | 1 | 123 tok/s (solo path) | 120 tok/s | | 2 | 57 tok/s | 113 tok/s | | 4 | 51 tok/s | 205 tok/s | | 8 | 45 tok/s | 360 tok/s | Without the pipeline, 8 slots ran at ~180 tok/s; without `--trim-stage-weights`, ~105 tok/s (the cards held 76-85 % of the experts instead of 84-100 %, and the CPU computed the rest for every row). ### Not done (yet) - MTP drafts in batch windows are argmax drafts (sampled rows accepted when the pick equals the draft). - Penalties are not applied in batch windows. - `--batch-groups` needs each stage on its own GPU; the slot sessions take VRAM and pinned RAM like the stage's session. - Not wired into `setup.py`: `tools/autoconfig.py` writes them into a copy of the config. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Mehr auf der Site
Links zu Install, Modellen, Releases.