Pull requests / #205

#205 fix(engine): clear stale error/cancellation string between requests and before prompt prefill

closed · @praveshkhatana · 0 comentarios · En GitHub

Server & APIModels & quants

Descripción

## Summary

In v0.1.26 and v0.1.27, if any request is cancelled while reading prompt chunks (e.g. client disconnects or times out on long contexts), the variable `err` is set to `"cancelled"`. 

Because `std::string err` is declared in `int main()` (around line 1311 of `src/program/generate.cpp`) and is never cleared at the start of a request, this string lingers into subsequent requests.

In commit `c6c65947` (E-9), `sp.on_chunk` was changed to:
```cpp
const bool batched = !multi_gpu && sp.draft_kv(mtp, R_rows, nxt.data(), T, p0, e);
if (!e.empty() || (!batched && !mtp.prefill(R_rows, nxt.data(), T, p0, e))) return false;
```
Because `e` is passed by reference from `Prefill::run(..., err)`, `!e.empty()` evaluates to `true` on chunk 0 of the *next* request. This causes `sp.on_chunk` to immediately abort without executing, `Prefill::run` returns `false`, and `generate.cpp` hits:
```cpp
if (!sp_ok) {
    if (!stop_req.load()) {
        std::fprintf(stderr, "strata serve: %s\n", err.c_str());
        std::printf("ERR %s\n", err.c_str());
        return 1;
    }
...
```
Because the current request was not stopped (`stop_req` is false), the engine outputs `strata serve: cancelled`, sends `ERR cancelled` to stdout, and terminates with exit code 1.

---

## Reproduction Steps

1. Start `strata --serve` on v0.1.26 or v0.1.27.
2. Send a request with a large context and abort/disconnect the client midway through prompt processing (triggers `STOP` -> `err = "cancelled"`).
3. Send a new valid request that exceeds `--short-read` (e.g. prompt length > 64 tokens, requiring batched prefill).
4. **Result:** The engine aborts on chunk 0, prints `strata serve: cancelled`, and exits with code 1.

---

## Fix

1. Call `err.clear()` right after `stop_req.store(false)` in `while (next_line(line))` in `src/program/generate.cpp`.
2. Call `err.clear()` before iterating prompt segments in `for (const int64_t to : {reread_to, root_at, turn_at, n - 1})`.
3. Call `err.clear()` on entry to `Prefill::run(...)` in `src/prefill/prefill.cpp`.

En el sitio

Enlaces a install, modelos, releases.