Pull requests / #637

#637 serve: retry an engine start that fails, and keep the last known context after a failed restart

closed · @JeanP00l · 0 comentarios · En GitHub

Server & APIMulti-GPUModels & quantsLinux

Descripción

## What

After the engine was ended mid-request (seen in long Claude Code sessions, many turns at ~20K tokens, on 2x 16 GB GPUs), the new engine could exit before `READY`: it ran out of VRAM while the old one's was still being handed back. `restart()` raised, `max_context` stayed 0, and from then on every request failed with `400 ... exceeds the context (0)` until the server was restarted by hand.

Since 0.1.38, `close()` already ends the old process and waits for it. This PR adds only what is still missing.

## Change (serve/server.py)

- A start that exits before `READY` is tried again, 3 times, `RESTART_RETRY_S` (15 s) apart. `starting` is set meanwhile, so a request gets the 503 "starting" of #344.
- `known_ctx` keeps the last context the engine reported. After a restart that failed for good, a request is checked against it and reaches `run()`, which starts the engine again, instead of getting a 400 about context 0.

## Tests

Two new tests in `RestartWindow`, both fail without the change:
- a start that fails once is retried and the engine comes up (fake engine script, retry delay patched to 0);
- a request after a failed restart (`max_context` 0, not starting, `known_ctx` 4096) gets 200, not a 400.

`python -m unittest serve.test_server`: 120 tests OK (Linux, Python 3.12).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

En el sitio

Enlaces a install, modelos, releases.