Pull requests / #18

#18 serve: treat unset/-1 max_tokens as remaining context, not 1024

closed · merged 2026-09-27 · @coolio986 · 0 comments · View on GitHub

Server & APIModels & quants

Description

Clients that send `max_tokens: -1` (unlimited), `0`, `null`, or omit it were capped at 1024 tokens — the old `or 1024` fallback, plus the `<= 0 -> 1024` clamp added in #7. Replies stopped at 1024.

## Changes
- `Service.prepare(..., max_new=None)` resolves the budget after tokenizing and returns `(ids, thinking, max_new)`. Unset/non-positive budget = remaining context.
- `CTX_SLACK = 8`: `strata --serve` rejects `prompt + max_new + 8 > context` (`src/program/generate.cpp`), so Python now uses the same margin. Unlimited = `max_context - 8 - len(ids)`; explicit budgets are checked against the same room (400 up front instead of an engine `ERR` mid-stream).
- Prompt that already fills the context with an unlimited budget -> 400 "leaves no room to answer".
- Explicit positive `max_tokens` / `max_completion_tokens` still honored exactly; `_debug_req` logs the resolved budget.

## Tests
New `serve/test_server.py` (unittest, MockEngine behind a real HTTP server): -1/0/missing/null on both APIs, explicit budgets, `max_completion_tokens` precedence, explicit overflow -> 400, near-full prompt.

`python -m unittest serve.test_server` -> 5 tests OK. Reverting to the 1024 fallback fails 14–16 checks.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.