Pull requests / #1288

#1288 serve: a long read gives way by what is left to read, not prompt length (#656)

open · @Jackwwg83 · 0 コメント · GitHub で見る

Server & APINVIDIA / CUDADocumentation

本文

#656 lets a long prompt read give way at a chunk boundary to a waiting request whose prompt is under half as long. It compares prompt lengths only. An agent's requests all carry a long system prompt and the history, so a waiting follow-up turn is almost never under half of a fresh read: it waits for the whole read, though it needs only a few hundred tokens. In a replay of a multi-user agent workload (a 27K-token system prompt, 5-15 users, RTX PRO 5000), 2,086 prompt reads gave way 1 time.

### What changes (serve/server.py only)
The rule compares tokens left to read on both sides:
- a waiting request: its prompt less the start the engine holds for sure - all a slot holds, or a slot's or an earlier read's prompt up to its last turn token (`<|im_start|>`, where the engine checkpoints a prompt and the next turn resumes), while a slot still holds it or the conversation cache parked it (only its slots' worth of the latest reads);
- the read on the control lines: what its PP lines say is left (reset when a read is sent; a request that arrives no longer clears it).

A read gives way only while the waiting request gets a slot at once (an admission: one free slot; a solo read: two, one for its parked part), so a request never waits on the control lines for a slot that a parked part needs to go on. Only prompts read to their end are noted. Each slot's held list is packed once; whole prefixes compare with `startswith` (8 slots of 50K tokens and 16 waiting prompts: under 50 ms a check).

A wrong guess (an evicted conversation, for example) costs one give-way, and a request gives way at most `YIELDS_MAX` (2) times; outputs do not change (`batch_interleave_test` covers a read that gives way).

Behaviour change: a fresh short prompt no longer makes a long read give way when the read has less than twice that left (the old rule compared full lengths).

### Tests
- `test_a_long_read_gives_way_to_a_follow_up_its_slot_holds`: a follow-up whose conversation a slot holds is answered before a fresh read of about the same prompt length. On main it fails (2.25 s against 1.61 s).
- `GiveWayByWhatIsLeft`: the rule case by case (partial prefixes, reads noted only at their end, the turn-token resume point, the free-slot condition, the read's own progress) and its cost.
- The server suite passes (457).

### Measured (RTX PRO 5000 48 GB, the 0.1.40 engine, only server.py differs)
`"parallel": 8` with the conversation cache on. Public text (the repository's own code and docs). B's first turn (~20K tokens) is answered; a fresh ~30K-token prompt A starts its read; 1 s later B's second turn arrives (about 24 new tokens). Median of 3 rounds:

| | 0.1.40 | with this change |
|---|---|---|
| B's next turn, time to first token | 6.65 s | 3.54 s |
| A (fresh), time to first token | 6.99 s | 7.80 s |
| reads that gave way | 0 of 3 | 2 of 3 |

The engine takes the give-way at the chunk boundary after the one where the server asks (8,192-token chunks here), and not when less than a chunk is left (the third round).

In our own 8-user replay (95% of prompt tokens reused, so 276 of 325 reads needed under 2K new tokens), the new rule gave way 1 time; the time to first token was 1.54 s median / 4.43 s p90 with the old rule and 1.35 s / 4.63 s with the new one (the p90 within noise). The long reads there were mostly two chunks (12-16K new tokens): after the first boundary one chunk or less is left, and the engine does not give way then. So the change matters where a read of three or more chunks meets a cached follow-up.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。