Pull requests / #7

#7 serve: fix zero-token replies from engine-queue race between requests

closed · merged 2026-09-26 · @coolio986 · 0 评论 · 在 GitHub 查看

BenchmarksServer & APIModels & quants

描述

## Summary

Fixes blank/empty responses that appeared when a client (e.g. opencode) has two requests in flight at once — a title-generation request alongside the main tool-calling turn. The second request frequently returned zero tokens: the server logged `done: 0 tokens in 0 s (length, cancel=False)` and the client showed an empty reply.

### Root cause

`StrataEngine` multiplexes every request over one resident `strata --serve` subprocess and a single shared line queue (`self.lines`), with no per-request framing. In `Service.run()`, a stop token (`<|im_end|>`) breaks the token loop with the engine generator still suspended and its internal `done` flag `False`. That generator's cleanup — write `STOP`, then drain `self.lines` until `DONE` — therefore ran only when `run()`'s frame was torn down later, **after** the `with self.fifo` lock had been released. By then the next request held the lock and had issued its own `GEN`; the prior request's drain raced it on the shared queue and consumed the **new** request's `DONE`. The new request saw an immediate end-of-stream: 0 tokens, `finish` left at its default `"length"`.

### Fix

Close the engine generator **inside** the `with self.fifo` block via `try/finally`, so a request's `STOP`+drain completes while it still holds the lock and cannot leave stray lines on the shared queue for the next request to misread as its own `DONE`.

Also included:
- Clamp a non-positive `max_new` to the default (1024) in both `/v1/chat/completions` and `/v1/messages`. Some clients send `max_tokens: -1` ("unlimited"); that value passed straight through and `GEN -1` made the engine yield 0 tokens.
- Distinguish a client disconnect: catch `GeneratorExit` and set `finish="disconnect"` instead of leaving it as `"length"`.
- Diagnostics gated on `STRATA_DEBUG` (off by default): a per-request `_debug_req` line and a `raw:` dump of decoded model output at completion.
- Progress reporting: add tok/s and cancel state to the `done:` line; emit the progress line every 1 s instead of 15 s.

## Test plan

- [x] `python -c "import ast; ast.parse(open('serve/server.py').read())"` passes.
- [x] Mock engine (`--engine mock`): 2 concurrent streaming + sequential streaming + non-stream OpenAI + Anthropic requests all return content with correct finish reason; no `0 tokens (length)` regressions.
- [x] Real engine (`STRATA_DEBUG=1 ./run-iq3_xxs.sh`): full opencode agentic loop (title-gen, tool calls, file write/edit, test run, final answer) completes with every turn producing tokens; the previously-blank `tools=11` main turn now streams normally.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。