Pull requests / #194

#194 serve: clear the error string at the start of every request (fixes #183)

closed · @demon851113 · 0 Kommentare · Auf GitHub

Setup & installServer & APINVIDIA / CUDAModels & quants

Beschreibung

Fixes #183.

### Root cause

`err` in `main()` is one string shared by every `--serve` request, and nothing clears it between requests. When a request is cancelled while it reads the prompt, `Prefill::run` sets `err = "cancelled"`. The serve loop treats that as a cancel rather than a failure, because `stop_req` is set, and finishes with `DONE ... cancel`, but it leaves `err` holding `"cancelled"`.

The next request clears `stop_req` but not `err`. When it reads its prompt through the batched path, `sp.on_chunk` starts with

```cpp
if (!e.empty() || (!batched && !mtp.prefill(...))) return false;
```

so the stale `"cancelled"` string counts as a failure. `sp.run` returns false and `stop_req` is now false, so the loop prints `strata serve: cancelled` / `ERR cancelled` and `return 1`s. The server then reloads the whole model.

That matches the log in #183: the cancelled request ends cleanly, and the *next* request fails with the misleading `cancelled` error, a few seconds in. A short next request goes through the window path, which does not check `err` up front, so it survives. That is why the crash only shows up when the next prompt is long.

### Fix

Clear `err` at the start of every request, next to the existing `stop_req.store(false)`. This is two lines, one of them a comment.

### Validation

RTX 5070 Ti 16 GB (sm_120), Ubuntu 24.04, CUDA 13.2, IQ3_XXS native pack, `--kv int8 --kv-resident 32768` (streamed KV, the same KV setup as the report), `--spec 4` + MTP. The engines were private (`serve.server.StrataEngine`, the class the HTTP server uses), with `--adapt-swaps 0`, greedy, and a 34,719-token prompt. The sequence was R1 = the prompt, cancelled 3 s in (a `cancel` event, which makes the wrapper send `STOP` exactly as on a client disconnect); R2 = the same prompt, 64 tokens; R3 = a short prompt.

| Engine | R1 | R2 (same prompt) | R3 (short) |
|---|---|---|---|
| `a790805` (0.1.27) | cancel | **`ERR cancelled`**, engine exits 1 | engine dead |
| `a790805` + this fix | cancel | OK: 64 tokens, resumes at the 16,384 checkpoint | OK |
| `a790805` + this fix, no R1 (fresh read) | – | OK: 64 tokens, **the same token ids** as the row above | OK |
| PR #189 + this fix, `--conversation-cache-mib 8192` | cancel | OK: the same 64 token ids | OK |

The unfixed engine also crashes through the HTTP server: a streaming `/v1/chat/completions` request disconnected 5 s into a 34,759-token prompt, and the next request got `the engine stopped unexpectedly (exit code 1)`.

The fix also means a checkpoint taken while the cancelled request was reading (here at 16,384) keeps working for the retry, as #183 asks under "Ideally". The resumed answer matches a fresh read token for token.

Mehr auf der Site

Links zu Install, Modellen, Releases.