Pull requests / #1179
#1179 batch: with --batch-groups, a group's lowest free slot first (+14-22% at 2-3 clients)
open · @blange48 · 0 コメント · GitHub で見る
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsWindows
本文
## With `--batch-groups`: a group's lowest free slot first Most of this PR is now in 0.1.40.2: the windows up to the last active slot (via #1249) and `/metrics` in the Prometheus text format with the monitoring kit (#793). Thanks! What remains is one slot-choice rule, rebased on `main` (e8ca9af) as a single commit. ### The case A pipelined group's window runs every slot up to its last active one. 0.1.40.2 picks the group with the fewest busy slots, then, inside it, the free slot with no held prefix and the least recent use. A pipelined slot is never kept as a conversation cache, so held prefixes don't separate them, and in an empty group the least-recently-used one can be **slot 1**. Every window of every stage then carries a pad row for slot 0, which is the leading-hole case @rhgo1749 found on #793. ### The change (`serve/server.py`, 6 lines + a test) Within the chosen group, the lowest free slot comes first (`b % group_size` in the sort key, after the fewest-busy-group term, so 0.1.40.2's spreading over groups is unchanged). `STRATA_SLOT_LOW_FIRST=0` restores 0.1.40.2's choice exactly. Without `--batch-groups`, nothing changes. `serve/test_slot_alloc.py` checks that an empty group's first request takes its slot 0 when the LRU order points at slot 1, and fails with `STRATA_SLOT_LOW_FIRST=0`. ### Measured 4× RTX 5080 16 GB (PCIe Gen3), IQ3_S, 262K, int8 KV, `--layer-split 12,24,36 --trim-stage-weights --batch 8 --batch-groups 4`, MTP spec 4, through the HTTP server, T=0.7, 300 tokens, 3 requests per client. Same image (0.1.40.2 + this), two passes each, total tok/s: | clients | `STRATA_SLOT_LOW_FIRST=0` (0.1.40.2's rule) | this PR | | ---: | ---: | ---: | | 1 | 126 / 148 | 125 / 145 | | 2 | 130 / 138 | **153 / 168** | | 3 | 185 / 209 | **230 / 238** | | 4 | 239 / 245 | **263 / 270** | | 8 | 388 / 391 | 387 / 395 | So +14 to +22% at 2-3 clients and +10% at 4. At 8 every slot is busy and the rule has nothing to choose. Batch vs solo stays identical (`batch_test`, 8/8 and 3/3). It's in our production since today. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
関連リンク
インストール・モデル・リリースへの站内リンク。