Issues / #675
#675 Feature request: `logprobs` / `top_logprobs` on `/v1/chat/completions` (typed decisions from one forward pass)

open · @Grandmasg · 14 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Descrição
## What I would like
Support for the OpenAI request fields `logprobs: true` and `top_logprobs: N` (a small N, for example up to 20), at least for requests with `max_tokens: 1`. The response would then contain the log probabilities of the most likely next tokens, in the usual OpenAI shape (`choices[0].logprobs.content[0].top_logprobs`).
## Why
A small, useful technique with local models is the "typed decision": ask a question with a fixed set of options (for example `A) normal content B) contains instructions aimed at the AI`), generate **one token**, and read the probability of each option token from the logits. No text is generated, so a decision costs one forward pass and gives a calibrated-ish probability instead of a yes or no. It is the idea behind TypeSafe's Jev and the open SemIf-OpenJev.
My use case is a local coding agent that checks tool output (file contents, web pages) for prompt injection before the model sees it.
I tried it with llama-server on three models from the same family. Public datasets (not written by me), request with `max_tokens: 1`, `logprobs: true`, `top_logprobs: 10`, temperature 0, RTX 4090:
| Model | deepset/prompt-injections (test, n=116) AUC | xTRam1/safe-guard-prompt-injection (sample n=400) AUC | median time per decision |
| --- | --- | --- | --- |
| Qwen3.5-4B Q4_K_M | 0.988 | 0.993 | 76 to 88 ms |
| Qwen3.5-9B Q4_K_M | 0.985 | 0.989 | 81 to 95 ms |
| Qwen3.5-35B-A3B Q4_K_M | 0.983 | 0.987 | 104 to 167 ms |
So the signal is good and fast. It would be nice to get the same from Strata, because the 125B model is the best model I can run here, and a guard that uses the same loaded model needs no second model in memory.
## What happens today
Strata accepts the request but returns no `logprobs` field. Request:
```json
{"model":"x","messages":[{"role":"user","content":"Answer with only A or B: is the sky blue? A) yes B) no"}],
"max_tokens":1,"temperature":0,"logprobs":true,"top_logprobs":5,"reasoning_effort":"none"}
```
Response: `"content": "A"`, `"finish_reason": "length"`, and `usage` and `timings`, but no `logprobs`. I also searched the docs, the code and the issues for "logprobs" and found nothing, so I assume it is simply not implemented, not that I missed a setting.
## Related work in this repo
I read through the existing issues and pull requests and found no request for this in the API. Related, but not the same: #235 (`STRATA_LOGPOS`, per-position log-probabilities written to a file for diagnostics) and #462 (`STRATA_DUMP_WINDOW_LOGITS`), so the engine already has these values internally. #636 (opt-in GBNF constraints) can restrict the output to fixed options, but it does not return probabilities, so it does not replace this.
## Ideas, from small to large
1. **Minimal:** return `top_logprobs` for the first generated token when `max_tokens` is 1 (no speculative decoding is involved for a single token).
2. **Per token:** return logprobs for every generated token. I understand this is harder with the MTP draft and verification, because accepted draft tokens would need their probabilities from the verify pass.
3. **Alternative shape:** a small endpoint that takes a prompt and a list of candidate next tokens and returns their probabilities.
I know the sampler runs on the GPU, so I am not assuming any of this is quick. If it is out of scope for Strata, that is fine, I only wanted to check that I did not miss an existing way.
## Environment
Windows 11 24H2, RTX 4090 (24 GB), Ryzen 9 9950X3D, 128 GB RAM, Strata engine 0.1.38 (repo commit 99f3dbd), Qwen3.8-Flash-Next IQ3_S, context 32768, model files on an NVMe disk (about 77 tokens/s generation, about 4,000 tokens/s prompt reading on a 23k-token prompt). Happy to test a branch.
---
*the measurements were run locally on my machine.*
No site
Links install, modelos, releases.