Pull requests / #334
#334 serve: enforce structured JSON in the native sampler
closed · @KadoBOT · 0 Kommentare · Auf GitHub
Setup & installServer & APINVIDIA / CUDADocumentationWindows
Beschreibung
Clients requesting structured JSON can currently receive successful prose because the OpenAI-compatible endpoint ignores `response_format`. This enforces `json_object` and `json_schema` during native decoding, so a required field such as `voiceover` cannot be omitted from a completed strict response. ### Behavior - Compile and validate schemas before loading or generating. Strict schemas require object roots, all properties required and `additionalProperties: false`. Unsupported or unenforceable schemas return HTTP 400. Only local references are resolved. - Use llguidance to compute a request-local token mask and apply it to native CUDA logits before token selection. Every selected token advances the grammar. Preserve schema property order, reasoning followed by JSON, and literal control-tag text inside JSON strings. Install llguidance and jsonschema through existing setup and Docker paths. - Structured requests use single-token verify windows and skip speculative drafting. Ordinary requests retain the existing speculative path. Masks do not affect prompt prefill or subsequent requests, including after cancellation. - Deliver incremental JSON prefixes through normal Chat Completions SSE. Token limits return `finish_reason: "length"`; OpenAI SDK parse/stream helpers raise their normal `LengthFinishReasonError`. Independently validate normally completed bodies before the terminal success chunk. Backend invariant failures return HTTP 502 or an SSE error. - Advertise `native_token_mask`, constrained decoding and incremental streaming through `/v1/status`. Require native `grammar_mask_v1` explicitly; do not fall back to prompting on older binaries. Structured formats with tools/MCP remain explicitly unsupported. One generation per request: no retry, repair pass or default insertion. This covers Chat Completions response formats, not the Responses API or provider-specific refusal policies. Native grammar masks currently disable speculative decoding for structured requests, so those requests trade throughput for enforcement. ### Validation - Built the native Windows engine using MSVC 19.44 and CUDA 13.0, targeting sm_120. - Native `grammar_mask_test` proves greedy and sampled decoding cannot select forbidden high-probability tokens. Existing `sampler_parity --selftest` passed. - `python -m unittest discover -s serve -p "test_*.py"`: 87 tests, OK, 3 optional skips. Focused structured/grammar checks: 14 tests, OK. OpenAI SDK checks use a local server and require no cloud key. - Live Windows Qwen generation passed the installed H3 Story and Fast schemas. A Fast streaming prompt explicitly asked to omit `voiceover`, `overlap` and `delivery`; the completed dialogue included all required fields. SDK parse and streaming parse, JSON arithmetic, token-limit handling, invalid-schema 400 and subsequent plain decoding passed. - Additional live checks passed reasoning followed by constrained JSON, literal tool/think tags plus Unicode inside a JSON string, and disconnect/STOP followed by a fresh request. Each request generated once. Model unloaded after testing. - Python compilation and `git diff --check` passed. Model files, local hardware/network settings and the user's H3 node are excluded. Usage and limits: `docs/structured-output.md`.
Mehr auf der Site
Links zu Install, Modellen, Releases.