Pull requests / #1328
#1328 serve + engine: constrained decoding for JSON formats; accept tools together with json_schema
open · @ncphri · 0 コメント · GitHub で見る
Server & APINVIDIA / CUDAModels & quantsWindows
本文
A JSON response format with tools was refused with 400; Codex sends both together. The two rejects are removed: the schema validates only the final answer turn, a tool-call turn passes through, and prepare_format(with_tools=True) adds the call-the-tools-first sentence to the directive. In Responses' JSON mode, text before a tool call is emitted as a message item before the function_call, as in plain mode, so Codex's replay keeps the cache prefix. When the engine reports token_mask=1, the answer is generated under a token mask built from the schema with llguidance (serve/constrain.py, include/strata/core/token_mask.hpp): the thinking stays free, the mask starts after </think>, one <tool_call> ends the turn's mask. MQ/MF/MK lines on the engine protocol; masked windows take no drafts. Post-hoc validation stays as the safety net. Fallbacks unchanged: no llguidance, no token_mask=1, --batch slots, images, or a schema llguidance cannot compile. Tests: serve/test_constrain.py (new), fake_mask_engine.py; 45 tests OK without a GPU. Real hardware: Windows, CUDA 13.0, sm_89, Swift 1.5 IQ3_XXS driven by Codex with tools + json_schema.
関連リンク
インストール・モデル・リリースへの站内リンク。