贡献 / #1288
#1288 serve: a long read gives way by what is left to read, not prompt length (#656)
open · @Jackwwg83 · 0 评论 · 去 GitHub 看
Server & APINVIDIA / CUDADocumentation
说明
#656 lets a long prompt read give way at a chunk boundary to a waiting request whose prompt is under half as long. It compares prompt lengths only. An agent's requests all carry a long system prompt and the history, so a waiting follow-up turn is almost never under half of a fresh read: it waits for the whole read, though it needs only a few hundred tokens. In a replay of a multi-user agent workload (a 27K-token system prompt, 5-15 users, RTX PRO 5000), 2,086 prompt reads gave way 1 time. ### What changes (serve/server.py only) The rule compares tokens left to read on both sides: - a waiting request: its prompt less the start the engine holds for sure - all a slot holds, or a slot's or an earlier read's prompt up to its last turn token (`<|im_start|>`, where the engine checkpoints a prompt and the next turn resumes), while a slot still holds it or the conversation cache parked it (only its slots' worth of the latest reads); - the read on the control lines: what its PP lines say is left (reset when a read is sent; a request that arrives no longer clears it). A read gives way only while the waiting request gets a slot at once (an admission: one free slot; a solo read: two, one for its parked part), so a request never waits on the control lines for a slot that a parked part needs to go on. Only prompts read to their end are noted. Each slot's held list is packed once; whole prefixes compare with `startswith` (8 slots of 50K tokens and 16 waiting prompts: under 50 ms a check). A wrong guess (an evicted conversation, for example) costs one give-way, and a request gives way at most `YIELDS_MAX` (2) times; outputs do not change (`batch_interleave_test` covers a read that gives way). Behaviour change: a fresh short prompt no longer makes a long read give way when the read has less than twice that left (the old rule compared full lengths). ### Tests - `test_a_long_read_gives_way_to_a_follow_up_its_slot_holds`: a follow-up whose conversation a slot holds is answered before a fresh read of about the same prompt length. On main it fails (2.25 s against 1.61 s). - `GiveWayByWhatIsLeft`: the rule case by case (partial prefixes, reads noted only at their end, the turn-token resume point, the free-slot condition, the read's own progress) and its cost. - The server suite passes (457). ### Measured (RTX PRO 5000 48 GB, the 0.1.40 engine, only server.py differs) `"parallel": 8` with the conversation cache on. Public text (the repository's own code and docs). B's first turn (~20K tokens) is answered; a fresh ~30K-token prompt A starts its read; 1 s later B's second turn arrives (about 24 new tokens). Median of 3 rounds: | | 0.1.40 | with this change | |---|---|---| | B's next turn, time to first token | 6.65 s | 3.54 s | | A (fresh), time to first token | 6.99 s | 7.80 s | | reads that gave way | 0 of 3 | 2 of 3 | The engine takes the give-way at the chunk boundary after the one where the server asks (8,192-token chunks here), and not when less than a chunk is left (the third round). In our own 8-user replay (95% of prompt tokens reused, so 276 of 325 reads needed under 2K new tokens), the new rule gave way 1 time; the time to first token was 1.54 s median / 4.43 s p90 with the old rule and 1.35 s / 4.63 s with the new one (the p90 within noise). The long reads there were mostly two chunks (12-16K new tokens): after the first boundary one chunk or less is left, and the engine does not give way then. So the change matters where a read of three or more chunks meets a cached follow-up. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
本站相关内容
相关页面的快捷入口。