Pull requests / #656

#656 serve: cooperative prefill preemption at safe chunk boundaries (--prefill-preempt)

closed · @j-luwierski · 0 comments · View on GitHub

Server & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows

Description

## Problem

Serving is sequential by design: the FIFO holds a mutex for a request's entire lifecycle, so a short interactive request sent during a long prefill waits for the whole prefill plus decode.

For orchestrator + sub-agent workloads sharing one Strata server, this resulted in tens of seconds of head-of-line blocking.

**Measured example:** `40.5 s` of queue wait behind a `50K`-token prefill.

---

## Solution

This PR adds **cooperative prefill preemption at safe chunk boundaries**.

A long-running request **A** (for example, a `200K`-token prompt) can park at a completed prefill chunk boundary, allowing a short request **B** to run immediately. Once B finishes, A is restored and continues as if it had never been interrupted.

The execution model remains intentionally simple:

- still **one sequence at a time**
- no decode batching
- no concurrent slots

### Engine (`--prefill-preempt`, opt-in)

- `Prefill::should_suspend`
  - Checked once per **completed chunk**
  - Single-GPU path only
  - Suspension happens only after the chunk completes
  - The normal `run` tail still executes, which:
    - commits `ple_prev` for the boundary
    - drains the compute stream
    - drains the copy stream

- `SuspReq`
  - Stores a complete suspension snapshot:
    - the running half via the existing checkpoint machinery
    - positional state of every owned QSA layer:
      - KV blocks
      - pooled rows
      - spare key
      - `idx_block_pos`
    - drafter KV state
  - Admission is **copy-only**
  - If snapshot creation fails, request A remains live and continues reading normally

- Engine protocol:
  - `SUSPENDED <pos> <total> id=N`
  - `RESUME id=N`
  - `CANCEL id=N`
  - `YIELD`
    - boundary offer emitted by the server
    - the engine itself cannot see the server request queue
  - Output lines may carry an `id=N` suffix for demultiplexing
  - Requests without an ID behave exactly as before

- Preemption gates:
  - minimum park position
  - per-request park cap
  - snapshot RAM budget admission (`T18`)
  - single GPU only
  - layer split fails closed

### Server (`--prefill-preempt`, opt-in)

The FIFO lock is split into **ownership periods** instead of being held for the request's entire lifecycle:

```text
GEN -> SUSPENDED
RESUME -> next SUSPENDED / DONE
```

A parked request waits outside the FIFO lock and receives keep-alives while suspended.

Additional behavior:

- if the parked request's waiter disappears, the park is withdrawn in place
- if the engine dies while a request is parked, the request fails cleanly
- `EngineRequest` supports `id=N` suffix demultiplexing

New observability:

`/metrics`

- `prefill_preemptions`
- `prefill_resumes`
- `max_queue_wait_ms`

`/status`

- `parked_requests`

---

## Results

Test system:

- **GPU:** RTX 4070 Ti SUPER
- **Quant:** IQ3_XXS
- **VRAM:** 16 GB

| Metric | Feature off | Feature on |
|---|---:|---:|
| B queue wait (`A prefill = 50K`) | `40.5 s` | `2.9 s` (**-93%**) |
| A total wall time | `43.8 s` | `45.3 s` (**+3.4%**) |

The task targets were:

- **≥ 80% reduction** in queue wait
- **< 10% penalty** to the interrupted request

Both targets are met with margin.

---

## Verification

### Bit-exact parity

Tokens and the complete state fingerprint are identical between:

1. uninterrupted request A
2. request A after `park -> interim request -> restore`

Verified state includes:

- `gdn`
- `ple`
- `tails`
- `pooled`
- `kv`
- `dead`
- `pooled_full`
- `mtp`
- `ple_prev`

Tested suspension boundaries:

- `4096`
- `8192`
- `12288`

An interim request was executed between park and restore with decode depths:

- `0`
- `1`
- `32`

The drafter KV state is included in the parity check.

### Trace diagnostics

Added diagnostics include:

- `STATE_POINT`
- `VERIFY_POINT`
- `WINDOW`

These provide:

- per-phase state hashes
- per-row verification-window fingerprints
- position
- offered-T count
- accepted count

### GPU scenarios

**12 GPU scenarios: ALL PASS**

Coverage includes:

- boundary tests ×5
- park-only tests ×3:
  - early
  - mid
  - late
- KV streaming:
  - mode `1`
  - residency-map reset
- repeat-3:
  - three consecutive parks
- cancel while parked
- verification that decode never yields

### Scheduler tests

Fake-engine scheduler suite: **9 tests**

Coverage includes:

- no contention
- preemption when another request queues
- FIFO ordering
- parked-request cancellation
- queued-request cancellation
- engine death while parked
- shutdown without deadlock

### Existing test suites

- `ctest`: **63/65**
- Python: **137 OK**

The two failing `ctest` cases are environmental and reproduce on pristine `main`.

---

## Bugs Found and Fixed During Verification

The tests caught several issues during development:

- A stale `YIELD` could re-park a resumed request at its next boundary.
- `preempt_count` was saved before the increment, allowing `--prefill-preempt-max` to be bypassed.
- A resumed tail with length `<= --short-read` could switch to verifier windows differently from an uninterrupted read.
- A lend-size layout mismatch caused an `E-9` failure across buffer layouts.

All of these were fixed before the final validation runs.

---

## Limitations / Future Work

Current limitations:

- only **one parked request** at a time
- image requests never park
- long-queue fairness currently uses a bounded wait:
  - `--prefill-preempt-max-wait-s`
  - default: `30 s`
- this is not yet a full scheduler
- snapshot RAM is not enforced through a single hard allocation budget
  - admission estimates the required memory and refuses before the large allocation
  - snapshot vectors are still separate allocations
  - allocation failure can therefore still surface as `bad_alloc`

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.