Pull requests / #1182

#1182 serve: poll cancellation during quiet prefill waits

closed · @hulkbig · 0 Kommentare · Auf GitHub

Server & APIAMD / HIPNVIDIA / CUDA

Beschreibung

Refs #962

While the native process is quiet before its first prompt-progress line or token, the Python server can wait ten seconds before checking a client's cancellation event. Poll the pipe in waits of at most 0.5 seconds while retaining the ten-second HTTP heartbeat and existing STOP / DONE / BADM cleanup ownership.

The regression tests use real HTTP disconnects and a scripted native subprocess. They cover serial generation, batch solo generation and batch admission across the three API endpoints, streaming and non-streaming; the following request must receive its own output. Additional tests preserve heartbeat cadence and silence deadlines.

Validation on macOS arm64 / Python 3.12.12: `python -m unittest serve.test_quiet_cancel serve.test_control_cancel serve.test_parallel serve.test_restart_waiters -v` — 30 tests passed. `git diff HEAD^ HEAD --check` passed.

This addresses the server-side quiet wait only. Actual CUDA/HIP prefill interruption, model execution and Claude Code cancellation behavior have not been tested, so this does not claim to fully resolve #962.

Developed, tested and reviewed with OpenAI coding assistants.

Mehr auf der Site

Links zu Install, Modellen, Releases.