Pull requests / #1504
#1504 serve --batch: a layer split reads a prompt in STRATA_SPLIT_PIECE-token runs, decided at every boundary
open · @sanastasiou · 0 comments · View on GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quants
Description
## What On a layer split under `serve --batch`, a prompt read that starts while a batch slot is decoding is cut into single chunks with a decode window between each, because the stage pipeline only overlaps within one `sp.run`. The read kept that shape to its last chunk, even when the slots finished their short turns a few seconds in. Long reads in a busy server ran at about half the speed of the same read on an idle server. `read_part` now decides at every chunk boundary how much to read next (`prompt_read_piece` / `prompt_read_run_end` in `include/strata/program/read_piece.hpp`): - a slot is decoding, or the engine is not a layer split: one chunk, as before; - layer split and no slot decoding: `STRATA_SPLIT_PIECE` tokens, rounded down to whole chunks (at least one); - `STRATA_SPLIT_PIECE` unset or 0: the whole remaining prompt in one run, as before. The piece size bounds how long a shorter request's `BYIELD` waits behind a long read. Scope: this applies with `--batch-groups 1` (the default). With `--batch-groups` > 1 on a split (`piped`), `read_part` still reads the whole prompt in one `sp.run`; this PR does not change that path. **Default behaviour does not change.** With `STRATA_SPLIT_PIECE` unset, an idle split reads the whole prompt in one run, the same as today. The only thing that differs without the variable is that a read which started beside a decoding slot goes back to one run once no slot decodes, where today it stayed on single chunks. ## Measured 2× RTX 3090 layer-split pair (PCIe 3.0, no NVLink), Qwen3.8-Flash-Next IQ3_S, `--batch 2 --prefill auto` (8K chunks), `STRATA_SPLIT_PIECE=65536`. Load: 4 parallel Claude Code compaction benchmarks, each growing to ~400K tokens, through one router onto two such pairs. | read | before | after | |---|---|---| | prompt reads that began beside a decoding slot | 1,240–1,740 tok/s | 1,800–2,450 tok/s | | e.g. a 124K read | 1,716 tok/s | 2,456 tok/s | | e.g. a 99.5K read at long context | 1,311–1,347 tok/s | 1,793 tok/s | | idle 249K read (unchanged path) | | 3,323 tok/s, 75 s | Over the whole 4×400K run: time to first byte p95 123 s → 76–80 s, mean wall time per conversation 1,260 s → 1,178 s. Recall 5/5 in all four conversations. ## Tests - `read_piece_test` (new, header-only, 15 checks): piece rules, rounding to whole chunks, run end clamped at the prompt end. - CUDA build: `generate.cpp` compiles; ran in production on both pairs. - `sycl/src/program/generate.cpp`: same edit, not compiled (no oneAPI here).
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.