Pull requests / #24
#24 serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting
closed · merged 2026-09-27 · @pjgmobile · 0 comentarios · En GitHub
Setup & installServer & APIModels & quantsSecurity
Descripción
# PR draft — Niko1221/Strata
**REBASED 2026-09-27 onto current main (89ba2dc, engine 0.1.8).** Branch `pr-fit-max-tokens`
lives in the worktree `/home/pete/strata-pr`; ready-to-send patch:
`/home/pete/pr-drafts/strata-fit-max-tokens.patch` (`git am strata-fit-max-tokens.patch`).
Your local working setup (main @ 9ef55a4 + local clamp patch) is untouched.
**Rebase notes:** all four previously-local upstream PRs (#7 zero-token, #8 conversation cache,
#9 cpu-pool sleep, #14 phantom-cores/dangling-else) are ALREADY in upstream main — the PR now
carries ONLY the clamp (11+/4- lines). Upstream's server.py evolved since the draft (CTX_SLACK
constant, 0/-1 = rest-of-context call semantics, progress lines): the resolution keeps all of
that and adds the flag. Conflicts were resolved in the worktree; mock-engine tests: strict
default = byte-identical 400, --fit-max-tokens = 200 with clamped generation.
**Title:** `serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting`
**Files:** `serve/server.py` (arg parsing + `Service.prepare()` + both call sites)
## Body
### Problem
Today the server rejects with 400 any request where `prompt_tokens + max_tokens` exceeds
`max_context` ("requests are never truncated"). That strictness mirrors llama.cpp and is the
right default for predictable single-purpose clients — but it breaks agentic clients that
legitimately ask for the model's full output cap with a long prompt
(e.g. prompt 131,888 + max_tokens 130,298 = 262,186 on a 262,144 window: rejected over a
42-token overshoot).
### Proposed change
Keep the current strict behaviour as the **default**, and add an opt-in flag:
```
--fit-max-tokens clamp max_tokens to the remaining context window instead of
rejecting the request (floor 16 tokens, 8-token safety headway)
```
In `Service.prepare()`:
```diff
--- a/serve/server.py
+++ b/serve/server.py
@@ def prepare(...)
ids, thinking = ...
- if len(ids) + requested_max > self.engine.max_context:
- raise ValueError(
- f"prompt ({len(ids)} tokens) + max tokens ({requested_max}) exceeds "
- f"the context ({self.engine.max_context}); requests are never truncated")
+ if len(ids) + requested_max > self.engine.max_context:
+ if not fit_max_tokens:
+ raise ValueError(
+ f"prompt ({len(ids)} tokens) + max tokens ({requested_max}) exceeds "
+ f"the context ({self.engine.max_context}); requests are never truncated")
+ room = self.engine.max_context - len(ids) - 8
+ if room <= 0:
+ raise ValueError(f"prompt ({len(ids)} tokens) fills the entire context "
+ f"({self.engine.max_context}); nothing left to generate")
+ requested_max = max(16, room)
```
(plus: thread the new flag from `main()` through config to `Service`, and update the two
call sites to use the returned clamped value — full working diff available, running in
production since 2026-09-26.)
### Verification
Prompt of 251,455 tokens + `max_tokens: 131072` on a 262,144 window, flag on:
**HTTP 200**, `completion_tokens: 10,681` (= 262,144 − 251,455 − 8), `finish_reason: length`,
`total_tokens: 262,136`. Flag off: the current 400, byte-identical. Small requests unaffected.
### Disclosure
*This PR was authored end-to-end by **GLM-5.3-Flash** (Z.ai): root-cause investigation, patch
authoring, and verification. Submitted after human review by the opener.*
En el sitio
Enlaces a install, modelos, releases.