Pull requests / #24

#24 serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting

closed · merged 2026-09-27 · @pjgmobile · 0 comentarios · En GitHub

Setup & installServer & APIModels & quantsSecurity

Descripción

# PR draft — Niko1221/Strata

**REBASED 2026-09-27 onto current main (89ba2dc, engine 0.1.8).** Branch `pr-fit-max-tokens`
lives in the worktree `/home/pete/strata-pr`; ready-to-send patch:
`/home/pete/pr-drafts/strata-fit-max-tokens.patch` (`git am strata-fit-max-tokens.patch`).
Your local working setup (main @ 9ef55a4 + local clamp patch) is untouched.

**Rebase notes:** all four previously-local upstream PRs (#7 zero-token, #8 conversation cache,
#9 cpu-pool sleep, #14 phantom-cores/dangling-else) are ALREADY in upstream main — the PR now
carries ONLY the clamp (11+/4- lines). Upstream's server.py evolved since the draft (CTX_SLACK
constant, 0/-1 = rest-of-context call semantics, progress lines): the resolution keeps all of
that and adds the flag. Conflicts were resolved in the worktree; mock-engine tests: strict
default = byte-identical 400, --fit-max-tokens = 200 with clamped generation.

**Title:** `serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting`

**Files:** `serve/server.py` (arg parsing + `Service.prepare()` + both call sites)

## Body

### Problem

Today the server rejects with 400 any request where `prompt_tokens + max_tokens` exceeds
`max_context` ("requests are never truncated"). That strictness mirrors llama.cpp and is the
right default for predictable single-purpose clients — but it breaks agentic clients that
legitimately ask for the model's full output cap with a long prompt
(e.g. prompt 131,888 + max_tokens 130,298 = 262,186 on a 262,144 window: rejected over a
42-token overshoot).

### Proposed change

Keep the current strict behaviour as the **default**, and add an opt-in flag:

```
--fit-max-tokens    clamp max_tokens to the remaining context window instead of
                    rejecting the request (floor 16 tokens, 8-token safety headway)
```

In `Service.prepare()`:

```diff
--- a/serve/server.py
+++ b/serve/server.py
@@ def prepare(...)
     ids, thinking = ...
-    if len(ids) + requested_max > self.engine.max_context:
-        raise ValueError(
-            f"prompt ({len(ids)} tokens) + max tokens ({requested_max}) exceeds "
-            f"the context ({self.engine.max_context}); requests are never truncated")
+    if len(ids) + requested_max > self.engine.max_context:
+        if not fit_max_tokens:
+            raise ValueError(
+                f"prompt ({len(ids)} tokens) + max tokens ({requested_max}) exceeds "
+                f"the context ({self.engine.max_context}); requests are never truncated")
+        room = self.engine.max_context - len(ids) - 8
+        if room <= 0:
+            raise ValueError(f"prompt ({len(ids)} tokens) fills the entire context "
+                             f"({self.engine.max_context}); nothing left to generate")
+        requested_max = max(16, room)
```

(plus: thread the new flag from `main()` through config to `Service`, and update the two
call sites to use the returned clamped value — full working diff available, running in
production since 2026-09-26.)

### Verification

Prompt of 251,455 tokens + `max_tokens: 131072` on a 262,144 window, flag on:
**HTTP 200**, `completion_tokens: 10,681` (= 262,144 − 251,455 − 8), `finish_reason: length`,
`total_tokens: 262,136`. Flag off: the current 400, byte-identical. Small requests unaffected.

### Disclosure

*This PR was authored end-to-end by **GLM-5.3-Flash** (Z.ai): root-cause investigation, patch
authoring, and verification. Submitted after human review by the opener.*

En el sitio

Enlaces a install, modelos, releases.