Pull requests / #1249
#1249 Pipelined batch: two requests run at half speed (pad rows in group windows, unbalanced slot choice)
closed · @crazyaimachine · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APINVIDIA / CUDAWindows
Beschreibung
Two small fixes for `--batch N --batch-groups G` on a layer split. Two concurrent requests ran at about half the speed of the 4-request case per request, sometimes worse. ### 1. A group window holds only the slots up to its last active one (`generate.cpp`) Every pipelined group window ran all `GS` slots of its group. An idle slot ran a pad row through every layer of every stage. The server spreads requests over the groups first (`slot_order`). So with `--batch 4 --batch-groups 2` and two requests, every window was half pad rows. Now `PGroup::S` is the group's last active slot + 1, and the window, its picks and `batch_launch()` cover only those slots. This fix was in the #559 series (ca74f92 there, "pipelined group windows only up to the last active slot"). It did not come along into 0.1.39, where the non-pipelined path got the same idea. ### 2. `pick_slot()` spreads requests over the groups (`serve/server.py`) When no slot held the prompt's start, `pick_slot()` took the least recently used free slot. Once every slot held an old conversation, two new requests could land in one group while the other group's stage sat idle. Now the rule is: 1. A held prefix of at least 512 tokens still wins, because reading a long conversation again costs more. 2. Otherwise the free slot in the group with the fewest busy slots is taken. 3. Then an empty slot. 4. Then the least recently used one. ### Measured Setup: - 2× TITAN RTX (Turing, PCIe 3.0), UD-Q4_K_XL with the BF16 PLE table. - `--batch 4 --batch-groups 2 --trim-stage-weights`, layer split 22, images on. - Through the HTTP server, T = 0.7, 400 tokens per answer, total tok/s. | Concurrent | 0.1.40.1 | with these two commits | | ---: | ---: | ---: | | 1 | 48–58 | 47–59 | | 2 | **36–38** (each request ~19 tok/s) | **52–82** | | 4 | 118–122 | 117–123 | Without `--batch-groups` (one group), 0.1.40.1 gave 50 / 68 / 94 on the same machine. ### Not solved Even with both fixes, a round of two requests now and then still runs at ~19 tok/s each, with the requests in slots 0 and 2 (different groups). In six back-to-back rounds the worst was 52 tok/s total, but it does happen. I have not found the cause yet. If you have an idea where to look (doorbells? the stage hand-off?), I am happy to test. The two commits are on 0.1.40.1 and apply to current `main`. Written by my agent, measured on my machine.
Mehr auf der Site
Links zu Install, Modellen, Releases.