Pull requests / #1281
#1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
closed · @routhjim · 0 comments · View on GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
Under sampling, `STRATA_SPEC_COUPLED=1` makes the draft layer sample its guess with the target's chain and the target's Philox draw. Both sides then pick by walking one uniform over their own candidates sorted by probability. When the draft's top_k / top_p list differs from the target's by even one token, every later interval shifts, and the two picks land on different tokens even where the distributions mostly agree. This PR adds an opt-in Gumbel-max pick that does not have that problem. With `STRATA_SPEC_GUMBEL=1` both the target and the coupled draft pick argmax_i p_i / E_i, where E_i ~ Exp(1) comes from a hash of (seed, counter, **token id**). That is exactly argmax(log p_i + Gumbel_i), so it is still an exact sample of p. A token gets the same noise in the draft's draw and in the target's, whichever other tokens survived either cut, so the two agree on the tokens they share. Verification is still exact-match against the target's pick: the text is a sample from the same distribution as before, and only the number of accepted drafts changes. **What changes** - `src/kernels/cuda/sampler.cu`: - `gumbel_exp(seed, counter, token)` (splitmix64 to Exp(1)). - The Gumbel pick in both tails: `sampler_kernel` and `sampled_tail_warp`, which serves the split, one-block and coupled-draft paths. - The coupled draft keys its noise by the real token id through `sub_to_id`, so the draft-vocab subset does not break the coupling. - `include/strata/kernels/sampler.hpp`: `SamplerParams::gumbel`. It sits in the struct's existing padding, so `sizeof` is unchanged (64 bytes). - `STRATA_SPEC_GUMBEL=1` is read once on the host. `sample_tokens` applies it to the target's pick, and `coupled_stage_kernel` writes it into the drafter's device copy of the params, so the captured round graphs see it. - **Off by default:** without the variable nothing changes, and greedy requests never read it. - `sampler_parity`: three new checks, run on all three sampled paths by the existing ctest entries: - the kernel's Gumbel pick equals a host reference draw for draw, under three chains (top_k/top_p, temperature only, and with min_p): 0 of 48 differ; - pick frequencies match the softmax over 10,240 draws (total variation 0.0098); - the flag is really read (2 of 16 sampled picks coincide with the inverse-CDF ones), and greedy output is unchanged. - Docs: a paragraph in `docs/DETAILS.md` next to the other opt-in switches, and a note in `include/strata/core/coupled_draft.hpp`. **Measured** on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, 128 GB, Linux 7.2, ROCm 7.14.1, the iGPU alone). - Setup: 82f46a8 plus this PR, built per `docs/STRIX_HALO.md` with the section-4 defaults and `STRATA_PF_FUSED_KQ=1`, Unsloth UD-Q4_K_XL, `--mmap-experts --expert-cache auto --spec 4 --mtp mtp/rt --lookup-chain 3 --kv int8` (the base model's draft layer, Q2_0 experts). - Requests: temperature 1.0, top_p 0.95, top_k 20, thinking off, a 1.3K-token prompt with a fresh prefix each time, 512 output tokens. - Counts come from the `/metrics` totals around each request; timings from `STRATA_DECODE_TIMING=1`. | arm | requests | drafts accepted | accepted tokens / request | tokens per window | ms per window (draft) | output tok/s | |---|---:|---:|---:|---:|---:|---:| | `STRATA_SPEC_COUPLED` off (default) | 12 | 52.8% | 318 | 2.65 | 63.6 (8.8) | 41.1 ± 0.8 | | `STRATA_SPEC_COUPLED=1 STRATA_SPEC_GUMBEL=1` | 13 | **59.9%** | **336** | **2.88** | 64.3 (9.1) | **44.9 ± 0.6** | A window costs the same: drafting by sampling adds about 0.3 ms of 64. The gain is entirely more tokens per verify window. ± is the standard error over the requests. Output stays an exact sample: with the switches on, three seeds give three different texts, and the same seed twice gives the same text. **Not tested:** - the CUDA build (I have no NVIDIA card), Windows, and other AMD cards; - `--batch-mtp` slot drafters; - `STRATA_SPEC_COUPLED=1` on its own with this draft layer: on this box with a different draft head it measured slightly below the default (one arm, within noise), so it is not claimed either way. Measured with the help of Claude (Anthropic). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.