Issues / #1666

#1666 Send a keep-alive while a tool call's array/object argument is being generated

open · @Efs-O · 0 comentários · No GitHub

BenchmarksServer & APIModels & quants

Descrição

**Problem**

With `stream_tools`, `OutputParser` streams string parameters piece by piece but holds every other type (arrays,
objects) until `</parameter>` (`serve/frontend.py`, the comment at `stream_tools`). The SSE `ping` is only sent when the
engine yields `None` (quiet) or during an MCP tool run. So while the model writes a large array argument — e.g. an
`edit_file` call whose `edits` is a list of 30-40 replacements — tokens are generated at full speed but nothing goes to
the client for minutes.

Clients with an idle timeout end the request. Forge's is 120 s; we hit it in a real agent turn: Strata generated
13.9K tokens, and the ~10K after the reasoning (the tool call) never reached the client.

**Reproduction** (v0.1.41, Qwen3.8 Flash-Next IQ3_S, 2 GPUs)

A streamed `/v1/chat/completions` with one tool whose argument is an array of objects, prompted to make 30 edits:

| | longest gap between SSE events | result |
|---|---|---|
| v0.1.41 | 27.8 s for a 7 KB array; minutes for a 66 KB one | client timeout on large calls |
| v0.1.41 + patch | 10.0 s | 7.6 KB array intact, `finish_reason: tool_calls` |

Meanwhile `/slots` `is_processing` and `/v1/status` `busy`/`generated` show the request is alive the whole time, so a
client *can* poll, but a keep-alive on the stream itself fixes every client at once.

**Suggested fix** (`serve/server.py`, +12 lines, patch below)

- `KEEPALIVE_S = 10.0`.
- In `Service.run`'s token loop, remember when anything was last sent (`last_out`: set on the heartbeat ping and when
  the parser emits events). When a token produces no event and nothing has gone out for `KEEPALIVE_S`, yield the same
  `"ping"` the heartbeat yields. No new event type; clients already ignore SSE comments.

**Tested**

- Full `serve` suite on v0.1.41 + the patch: 625 tests, 1 failure (`test_slots...test_relative_save_dir_becomes_absolute`),
  which fails identically on plain v0.1.41 on this machine (a TEMP 8.3 short-path issue, unrelated).
- Live: the reproduction above, and an 880 s Forge agent turn whose single `edit_file` call carried 66 KB of
  arguments, with no stall.

Not done: a unit test for the ping cadence (needs a fake generator yielding tokens with no events for >10 s; I can
add one with a patched clock if you want it).

**Patch** (against v0.1.41)

```diff
diff --git a/serve/server.py b/serve/server.py
index 77eee032..805236fc 100644
--- a/serve/server.py
+++ b/serve/server.py
@@ -196,6 +196,7 @@ def focused_recovery_prompt(tok, ids, generated):
 
 
 LOOPBACK_NAMES = ("localhost", "127.0.0.1", "::1")
+KEEPALIVE_S = 10.0          # see last_out in Service.run
 CTX_SLACK = 8               # `strata --serve` rejects prompt + max_new + 8 > context: keep the same margin here
 # The live tok/s is a rate over a window, not a mean since the first token: a mean reads ~1/elapsed at the first
 # token (the Monitor showed five-digit numbers) and then undershoots for the first second of every answer.
@@ -3508,9 +3509,15 @@ class Service:
                         seg, wrap, leaving = [], False, False   # this pass's tokens; the budget is reached; closed
                         opens = False                   # the thinking is over: write the forced call's opening
                         try:
+                            # A value the frontend holds back (a tool call's array or object
+                            # argument is sent whole, once complete) can keep the stream silent for minutes while
+                            # tokens are generated; a client's idle timeout then ends the request.  Send the same
+                            # keep-alive the engine's heartbeat sends when nothing has gone out for KEEPALIVE_S.
+                            last_out = time.perf_counter()
                             for t in gen:
                                 if t is None:               # heartbeat while the engine is quiet
                                     last_print = self._progress(last_print, st=st)
+                                    last_out = time.perf_counter()
                                     yield "ping", None
                                     continue
                                 n += 1
@@ -3540,6 +3547,11 @@ class Service:
                                     if ev.kind in ("content", "tool_start", "tool_call"):
                                         answered = True
                                     yield "event", ev
+                                if evs:
+                                    last_out = time.perf_counter()
+                                elif time.perf_counter() - last_out >= KEEPALIVE_S:
+                                    last_out = time.perf_counter()
+                                    yield "ping", None
                                 if stops is not None and stops.hit is not None:
                                     finish = "stop"         # gen.close() below STOPs the engine, as for a stop token
                                     break
```

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.