Pull requests / #636

#636 native: add opt-in GBNF constraints to existing generation

closed · @CC-David-CC · 0 コメント · GitHub で見る

Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

本文

this is selfish and not for you all. if you bring it in i dont have to worry about supporting it or losing it. i use gbnf not you. there's no power with gbnf. dont look. gbnf is not the function you are looking for *it totally is*
- Unconstrained: `The sky is purple and the number is 999.`
- With GBNF: `color=purple;count=999`
It does not include JSON Schema, Lark tools, or grammar combined with reasoning/tools.

Adds raw GBNF enforcement to Strata's existing native decoder, exposed through a `grammar` extension on Chat Completions and Responses. The native matcher masks token choices before the existing sampler and tracks the same retained tokens as model commit and emitted output. Requests fail explicitly when the build, model profile or grammar is unsupported.

This is a follow-up to the stateless Responses contribution associated with #451. The branch contains its R4 implementation. Land Responses first, then review this contribution against that landing; a comparison to today's upstream `main` still includes both features. GBNF is not required for Codex to speak Responses.

## Changes

- Persistent target-only `--serve --spec 1` uses the existing decoder without loading MTP weights. This establishes the reference mode without adding a second inference service.
- Optional pinned XGrammar supplies GBNF compilation and private matcher state. Dependency preparation is explicit; server startup does not download a backend. Compilation, source, output and matcher work have documented bounds.
- The HTTP service preflights the native `gbnf-v2` capability and tokenizer vocabulary before successful streaming headers. Both APIs forward the same constraint to the existing engine.
- Constrained MTP, coupled proposals and suffix lookup use tentative grammar state. One retained count bounds model commit, output and matcher advancement. There is no automatic unmasked retry or acceleration fallback.

The G0 target-only path and retained-window corrections also change native behavior outside the GBNF compile guard. They need review even when grammars are disabled; the compile flag does not isolate every change in this PR.

## Enablement

Native grammar support defaults off. Explicitly prepare the pinned dependencies and configure the existing native build with `-DSTRATA_ENABLE_GBNF=ON`. Start the existing server with its normal model settings. For Responses, separately install its optional requirements, set the replay key and enable `--experimental-responses`; Chat Completions does not need that API flag.

Send raw GBNF through `grammar` (SDK clients use `extra_body`). This is a Strata extension, not an OpenAI standard field. The [native guide](https://github.com/CC-David-CC/Strata-a5500/blob/c91260ccd3a929ae51696c9dff1f972d6d56adfe/docs/NATIVE_GBNF.md) gives build commands, target-only and speculative configurations, requests and limits.

## Recorded validation

- Linux/CUDA RTX 4090, Coder IQ1_M, pinned XGrammar and native build. G1-G5 reports retain exact commands and environments; no Windows native, HIP or multi-GPU qualification is claimed.
- G5 Python regression checkpoint: 253 passed, five optional skips; seven application-state tests. Native Release/sanitizer, parser/matcher, prefix-commit and GPU sampler checks are recorded separately.
- Fixed-logit GPU selection: 5,184 decodes across three sampler paths; seven GPU tests; CUDA memcheck reported zero errors. These isolate selection correctness rather than claiming identical model logits across graph shapes.
- Real native/API correctness and warm-performance matrices cover target-only, MTP, coupled and suffix configurations, plus ordinary unconstrained MTP/suffix regression cases. Numerical diagnostics retain observed model-logit differences. Narrow-grammar gains are workload-specific; slower Unicode cases remain in the evidence.
- A real pinned local Codex session completed a native `Hello world` request through Responses with external client HTTP blocked. This was unconstrained generation and did not exercise native tool calls.

See the [G5 report](https://github.com/CC-David-CC/Strata-a5500/blob/c91260ccd3a929ae51696c9dff1f972d6d56adfe/docs/gbnf-evidence/G5/REPORT.md) and [native Codex report](https://github.com/CC-David-CC/Strata-a5500/blob/c91260ccd3a929ae51696c9dff1f972d6d56adfe/docs/responses-evidence/native-codex/REPORT.md). Earlier synthetic protocol fixtures are labeled separately from native execution.

## Scope and remaining gates


- [ ] Land the Responses prerequisite and carry its accounting/replay follow-ups into this branch.
- [ ] Build and smoke-test the final native source with `STRATA_ENABLE_GBNF=OFF`. The existing off-build evidence is an earlier configuration check; final ordinary-generation regressions used the grammar-enabled build.
- [ ] Review the unconditional target-only and retained-window changes explicitly; split general native fixes if the maintainer prefers.
- [ ] Reduce repeated raw evidence in the upstream diff, retaining concise reports, necessary tests and links to the complete public archive.

Reviewed head: `c91260ccd3a929ae51696c9dff1f972d6d56adfe` on `CC-David-CC:work/gbnf`, starting from Responses implementation `96092670da0dc3c1cfcce99bb90a3e6ca25ae1d9`. The separate Responses tip's extra commit contains only its receipt; that receipt is also retained in this branch.

関連リンク

インストール・モデル・リリースへの站内リンク。