Pull requests / #334

#334 serve: enforce structured JSON in the native sampler

closed · @KadoBOT · 0 commentaires · Sur GitHub

Setup & installServer & APINVIDIA / CUDADocumentationWindows

Description

Clients requesting structured JSON can currently receive successful prose because the OpenAI-compatible endpoint ignores `response_format`. This enforces `json_object` and `json_schema` during native decoding, so a required field such as `voiceover` cannot be omitted from a completed strict response.

### Behavior

- Compile and validate schemas before loading or generating. Strict schemas require object roots, all properties required and `additionalProperties: false`. Unsupported or unenforceable schemas return HTTP 400. Only local references are resolved.
- Use llguidance to compute a request-local token mask and apply it to native CUDA logits before token selection. Every selected token advances the grammar. Preserve schema property order, reasoning followed by JSON, and literal control-tag text inside JSON strings. Install llguidance and jsonschema through existing setup and Docker paths.
- Structured requests use single-token verify windows and skip speculative drafting. Ordinary requests retain the existing speculative path. Masks do not affect prompt prefill or subsequent requests, including after cancellation.
- Deliver incremental JSON prefixes through normal Chat Completions SSE. Token limits return `finish_reason: "length"`; OpenAI SDK parse/stream helpers raise their normal `LengthFinishReasonError`. Independently validate normally completed bodies before the terminal success chunk. Backend invariant failures return HTTP 502 or an SSE error.
- Advertise `native_token_mask`, constrained decoding and incremental streaming through `/v1/status`. Require native `grammar_mask_v1` explicitly; do not fall back to prompting on older binaries. Structured formats with tools/MCP remain explicitly unsupported.

One generation per request: no retry, repair pass or default insertion. This covers Chat Completions response formats, not the Responses API or provider-specific refusal policies. Native grammar masks currently disable speculative decoding for structured requests, so those requests trade throughput for enforcement.

### Validation

- Built the native Windows engine using MSVC 19.44 and CUDA 13.0, targeting sm_120.
- Native `grammar_mask_test` proves greedy and sampled decoding cannot select forbidden high-probability tokens. Existing `sampler_parity --selftest` passed.
- `python -m unittest discover -s serve -p "test_*.py"`: 87 tests, OK, 3 optional skips. Focused structured/grammar checks: 14 tests, OK. OpenAI SDK checks use a local server and require no cloud key.
- Live Windows Qwen generation passed the installed H3 Story and Fast schemas. A Fast streaming prompt explicitly asked to omit `voiceover`, `overlap` and `delivery`; the completed dialogue included all required fields. SDK parse and streaming parse, JSON arithmetic, token-limit handling, invalid-schema 400 and subsequent plain decoding passed.
- Additional live checks passed reasoning followed by constrained JSON, literal tool/think tags plus Unicode inside a JSON string, and disconnect/STOP followed by a fresh request. Each request generated once. Model unloaded after testing.
- Python compilation and `git diff --check` passed. Model files, local hardware/network settings and the user's H3 node are excluded.

Usage and limits: `docs/structured-output.md`.

Sur le site

Liens install, modèles, releases.