Pull requests / #1646

#1646 fix(api): honor chat logit_bias on CUDA and HIP

open · @CC-David-CC · 0 comments · View on GitHub

Multi-GPUAMD / HIPNVIDIA / CUDADocumentation

Description

Resolves #1627.

Chat requests accepted logit_bias without applying it. This patch applies per-request biases in CUDA/HIP target sampling, including greedy decoding and ordinary MTP.

### Why Hello?

Hello is familiar, is one token (ID 9419) in the tested tokenizer, and is naturally emitted for this prompt. 9419 is a token ID, not a logit/probability; -100 is the requested bias and means exclusion here. Nothing is banned by default.

Same prompt: "Reply with one English greeting, one word only." Temperature 0, max_tokens 8, reasoning_effort "none". Only logit_bias changes:

| Request value | Measured output |
|---|---|
| Field omitted | Hello (9419) |
| `"logit_bias": {"9419": -100}` | Hi (12675) |
| `"logit_bias": [[9419, false]]` | Hi (12675) |
| Field omitted again | Hello (9419) |

All four English checks passed on llm-60. Token IDs depend on the tokenizer; a token ban is not a universal word ban.

52 further HTTP checks passed across CUDA and HIP, including MTP, streaming and 103,215 bans. Both builds, 276 Python tests per backend, parser tests and three GPU sampler parity paths per GPU passed. Matched no-bias outputs agree with pristine main fb58e0db.

Invalid input fails before streaming. Older engines, SYCL and continuous batching reject nonempty biases. Multi-GPU and experimental coupled/probabilistic MTP were not validated. No throughput claim.

[Full before/after requests and raw results](https://github.com/CC-David-CC/Strata-a5500/blob/fix/chat-logit-bias-1627/docs/LOGIT_BIAS.md)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.