Pull requests / #19

#19 serve: per-request temperature/top_p/top_k/min_p/penalties sampling

closed · merged 2026-09-27 · @j-luwierski · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

Description

## Problem

Temperature (and the rest of the sampling block) was ignored end to end:

- the verify window's head sampling was pinned to greedy INSIDE the captured CUDA graph (`verify.cpp` hardcoded `sp.greedy = true`) - and a captured sampler would replay the same draws forever;
- the GEN protocol carried no sampling keys;
- `StrataEngine.generate` accepted the request's sampling parameters and dropped them, so every served request decoded greedy whatever the client asked - the README's "temperature / top_p / top_k / seed are honored per request" was not true over the HTTP API.

## What changed

**Per-request temperature / top_p / top_k / min_p / seed, plus the llama.cpp penalty set.** The sampled chain's options are live end to end: an OpenAI-style request (`temperature`, `top_p`, `top_k`, `min_p`, `presence_penalty`, `frequency_penalty`, `repetition_penalty`, `penalty_last_n`, `seed`) or the config's `sampling` block reaches the engine's per-request sampler. Absent / neutral keys cost nothing and stay greedy, exactly as before.

**The sampling moves out of the captured graph.** The graph computes the head logits; `run()` applies `sample_tokens` host-side after the replay, with that request's own parameters and a fresh Philox counter per window. The GEN protocol accepts optional `key=value` pairs before the ids (`temperature=`, `top_p=`, `top_k=`, `min_p=`, `penalty_repeat=`, `penalty_freq=`, `penalty_present=`, `penalty_last_n=`, `seed=`); absent keys stay greedy. Emitted tokens are always the TARGET's own samples (the drafts only fill the verified window), so the output distribution is the target's regardless of what the drafts propose.

**A persistent per-request draw counter; sampled requests default top_k=20.** The sampled path requires top_k in 1..64 (the selection loop is O(k·n_vocab) by design); with the engine-default top_k=0 it returned without writing the token, leaking stale buffer contents as output - sampled requests now default to the sampler's own top_k=20. The Philox draw counter also persists across a request's verify windows (it was run-local before: every window redrew the same numbers, so identical requests returned identical sampled output).

**min_p and the penalties.** `min_p` cuts the top_k survivors to the prefix with p >= min_p · p_max - in logits, `sel_logit >= sel_logit[0] + logf(min_p)` - before top_p, so the two cuts compose; 0 disables, and the head always survives, so the kept count never reaches zero. The penalties (repeat/freq/present) count over a history window: the serve loop uploads the request's last `penalty_last_n` tokens - the consumed tail plus the fed-back head - into one device row before every window; a request without penalties hands the sampler a null row and takes the byte-for-byte neutral path. The counts come from a shared per-token bitmap: each block builds it from the history row in one `atomicOr` pass, and a candidate pays the O(hlen) count scan only when its bit is set - the naive per-candidate recount cost O(k·n_vocab·hlen) integer compares per token (~318M at top_k=20 on the 248,320 vocabulary) and dragged sampled decode 45.2 → 30.8 tok/s on a real chat workload. The repeat penalty multiplies for non-positive logits and divides for positive ones (llama.cpp's chain; dividing unconditionally inverts the penalty on half the vocabulary), and the presence penalty is a boolean, not the count.

**The sampled chain is one block per token.** It ran in ONE thread per token: top_k alone was k sequential full-vocabulary scans and a verify window measured ~1.6 s in the sampler - sampled requests decoded at 1.5 tok/s against 54 greedy. The selection is k argmax rounds, so it parallelises exactly like the greedy kernel: one block per token, the rounds back to back inside it, each a two-level (warp, then block) reduction. Ties resolve to the lowest index (the serial scan's strict `>`), so the kept sequence - set and order - is unchanged; the penalty, top_p cut, temperature and Philox draw read that order as before.

**server.py forwards the block to the engine.** `sampling_keys` forwards each field only when it is not neutral, so a client that leaves them out costs nothing; an absent temperature keeps the engine's greedy default and temperature=0 is not forwarded (it means the same thing); top_k outside the sampled path's 1..64 is dropped rather than sending a value the engine would refuse; a penalty without a window sends the engine's default `penalty_last_n=64`. The run config's optional `sampling` block sets the defaults for fields a request leaves out - the request's own fields always win, a null field falls back to the default, and a value the engine cannot honour refuses to start the server (a typo'd config should not quietly change sampling).

**Docs.** DETAILS.md: the sampling + reproducibility note (static residency), the config defaults, and the v1 limits line no longer claims temperature is ignored.

## Verification

- greedy requests stay bit-identical; sampled requests with the same seed replay identically (static residency)
- end to end over the HTTP server: temperature=0 and absent agree token for token; a temperature=0.8 request differs from greedy; two temperature=0.8 seed=99 requests return identical text
- `sampler_parity` passes, grown to a full-chain host reference (the Philox draw transcribed host-side) plus fixtures for the sampled chain with penalties, for min_p at 0/0.5/0.9, and for the penalty-window clamp - each fixture first asserts it can SEE the feature before asserting the kernel matches; a 248,320-vocab GPU-vs-CPU draw check matches the serial reference on five parameter sets including injected duplicate logits
- on the IQ3_XXS artifact, same seed: neutral penalty keys are token-identical to no keys; a six-copy prompt that loops 24/24 on one token clean produces 24 distinct tokens with repeat 1.5 + present 0.5; min_p 0.9 changes the pick
- sampled decode 1.50 → 42.7 tok/s (RTX 4070 TS, IQ3_XXS, 128-token generations); greedy decode unchanged (53.3 vs 54.3 tok/s median, within run noise - greedy launches the untouched kernel)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.