Issues / #1615
#1615 [Feature]: let harnesses recognize the context-overflow 400 (include a phrase they already match)
open · @donnielrt · 0 commentaires · Sur GitHub
Description
*Filed by an AI agent (Claude, running in Claude Code) on behalf of the repository operator, who reviewed and approved this submission.* ### What do you want to do? Let agent harnesses recognize Strata's context-overflow 400 as an overflow, so they run their own compaction-and-retry instead of showing a dead error. **What happens today.** `serve/server.py` (main, fb58e0d) refuses an oversized request with one of two messages: ``` prompt (N tokens) leaves no room to answer in the context (C); requests are never truncated prompt (N tokens) + max tokens (M) exceeds the context (C); requests are never truncated. Send a smaller max_tokens (at most R here), or add "fit_max_tokens": true to ... ``` as HTTP 400 with `"type": "invalid_request_error"`. Harnesses detect overflow by matching the error text against the phrases servers already use. Pi (pi-coding-agent, `pi-ai/utils/overflow.js`, `isContextOverflow`) matches, among others: - llama.cpp server: `exceeds the available context size` - OpenAI: `exceeds the context window` - generic: `context_length_exceeded` / `context length exceeded`, `too many tokens`, `token limit exceeded` Strata's wording matches none of them, so Pi treats the 400 as an ordinary provider error: no compaction, the turn fails, and the session is stuck until the user shrinks it by hand. We hit this on 0.1.40.1 with a 150,722-token prompt on a 131,072 context (one large tool result pushed the prompt past the limit in a single turn). `fit_max_tokens` cannot help in that case because the prompt alone exceeds the context (the `room < 1` branch). For comparison, llama.cpp's server returns `request (N tokens) exceeds the available context size (C tokens), try increasing it` with `"type": "exceed_context_size_error"` plus `n_prompt_tokens` and `n_ctx` fields, and Pi recovers from that automatically. **Suggestion.** Keep the current message and guidance, and add a phrase harnesses already match. The smallest change is appending llama.cpp's wording to both messages, e.g. ``` prompt (N tokens) + max tokens (M) exceeds the context (C); the request exceeds the available context size. Requests are never truncated. Send a smaller max_tokens ... ``` Optionally also return `"type": "exceed_context_size_error"` with `n_prompt_tokens` and `n_ctx`, matching llama.cpp, for clients that check the type instead of the text. **Workaround we use today.** Our Pi extension rewrites the error text on the client side to include "the request exceeds the available context size" when it matches Strata's two messages, which makes Pi's compact-and-retry fire. We'd rather delete that adapter once the server says it itself.
Sur le site
Liens install, modèles, releases.