Pull requests / #656
#656 serve: cooperative prefill preemption at safe chunk boundaries (--prefill-preempt)
closed · @j-luwierski · 0 评论 · 在 GitHub 查看
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
描述
## Problem
Serving is sequential by design: the FIFO holds a mutex for a request's entire lifecycle, so a short interactive request sent during a long prefill waits for the whole prefill plus decode.
For orchestrator + sub-agent workloads sharing one Strata server, this resulted in tens of seconds of head-of-line blocking.
**Measured example:** `40.5 s` of queue wait behind a `50K`-token prefill.
---
## Solution
This PR adds **cooperative prefill preemption at safe chunk boundaries**.
A long-running request **A** (for example, a `200K`-token prompt) can park at a completed prefill chunk boundary, allowing a short request **B** to run immediately. Once B finishes, A is restored and continues as if it had never been interrupted.
The execution model remains intentionally simple:
- still **one sequence at a time**
- no decode batching
- no concurrent slots
### Engine (`--prefill-preempt`, opt-in)
- `Prefill::should_suspend`
- Checked once per **completed chunk**
- Single-GPU path only
- Suspension happens only after the chunk completes
- The normal `run` tail still executes, which:
- commits `ple_prev` for the boundary
- drains the compute stream
- drains the copy stream
- `SuspReq`
- Stores a complete suspension snapshot:
- the running half via the existing checkpoint machinery
- positional state of every owned QSA layer:
- KV blocks
- pooled rows
- spare key
- `idx_block_pos`
- drafter KV state
- Admission is **copy-only**
- If snapshot creation fails, request A remains live and continues reading normally
- Engine protocol:
- `SUSPENDED <pos> <total> id=N`
- `RESUME id=N`
- `CANCEL id=N`
- `YIELD`
- boundary offer emitted by the server
- the engine itself cannot see the server request queue
- Output lines may carry an `id=N` suffix for demultiplexing
- Requests without an ID behave exactly as before
- Preemption gates:
- minimum park position
- per-request park cap
- snapshot RAM budget admission (`T18`)
- single GPU only
- layer split fails closed
### Server (`--prefill-preempt`, opt-in)
The FIFO lock is split into **ownership periods** instead of being held for the request's entire lifecycle:
```text
GEN -> SUSPENDED
RESUME -> next SUSPENDED / DONE
```
A parked request waits outside the FIFO lock and receives keep-alives while suspended.
Additional behavior:
- if the parked request's waiter disappears, the park is withdrawn in place
- if the engine dies while a request is parked, the request fails cleanly
- `EngineRequest` supports `id=N` suffix demultiplexing
New observability:
`/metrics`
- `prefill_preemptions`
- `prefill_resumes`
- `max_queue_wait_ms`
`/status`
- `parked_requests`
---
## Results
Test system:
- **GPU:** RTX 4070 Ti SUPER
- **Quant:** IQ3_XXS
- **VRAM:** 16 GB
| Metric | Feature off | Feature on |
|---|---:|---:|
| B queue wait (`A prefill = 50K`) | `40.5 s` | `2.9 s` (**-93%**) |
| A total wall time | `43.8 s` | `45.3 s` (**+3.4%**) |
The task targets were:
- **≥ 80% reduction** in queue wait
- **< 10% penalty** to the interrupted request
Both targets are met with margin.
---
## Verification
### Bit-exact parity
Tokens and the complete state fingerprint are identical between:
1. uninterrupted request A
2. request A after `park -> interim request -> restore`
Verified state includes:
- `gdn`
- `ple`
- `tails`
- `pooled`
- `kv`
- `dead`
- `pooled_full`
- `mtp`
- `ple_prev`
Tested suspension boundaries:
- `4096`
- `8192`
- `12288`
An interim request was executed between park and restore with decode depths:
- `0`
- `1`
- `32`
The drafter KV state is included in the parity check.
### Trace diagnostics
Added diagnostics include:
- `STATE_POINT`
- `VERIFY_POINT`
- `WINDOW`
These provide:
- per-phase state hashes
- per-row verification-window fingerprints
- position
- offered-T count
- accepted count
### GPU scenarios
**12 GPU scenarios: ALL PASS**
Coverage includes:
- boundary tests ×5
- park-only tests ×3:
- early
- mid
- late
- KV streaming:
- mode `1`
- residency-map reset
- repeat-3:
- three consecutive parks
- cancel while parked
- verification that decode never yields
### Scheduler tests
Fake-engine scheduler suite: **9 tests**
Coverage includes:
- no contention
- preemption when another request queues
- FIFO ordering
- parked-request cancellation
- queued-request cancellation
- engine death while parked
- shutdown without deadlock
### Existing test suites
- `ctest`: **63/65**
- Python: **137 OK**
The two failing `ctest` cases are environmental and reproduce on pristine `main`.
---
## Bugs Found and Fixed During Verification
The tests caught several issues during development:
- A stale `YIELD` could re-park a resumed request at its next boundary.
- `preempt_count` was saved before the increment, allowing `--prefill-preempt-max` to be bypassed.
- A resumed tail with length `<= --short-read` could switch to verifier windows differently from an uninterrupted read.
- A lend-size layout mismatch caused an `E-9` failure across buffer layouts.
All of these were fixed before the final validation runs.
---
## Limitations / Future Work
Current limitations:
- only **one parked request** at a time
- image requests never park
- long-queue fairness currently uses a bounded wait:
- `--prefill-preempt-max-wait-s`
- default: `30 s`
- this is not yet a full scheduler
- snapshot RAM is not enforced through a single hard allocation budget
- admission estimates the required memory and refuses before the large allocation
- snapshot vectors are still separate allocations
- allocation failure can therefore still surface as `bad_alloc`站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。