Pull requests / #59
#59 Fix/53 sampler penalties. sampler: penalties apply once, as in llama.cpp's default chain (#53); guard the unsized penalty bitmap
closed · @j-luwierski · 0 评论 · 在 GitHub 查看
Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
描述
Fixes #53.
## What happened
The block-per-token sampler rewrite (4d39f7a) added penalties to the `top_k` selection and kept the
old application in the final stage, so the sampled chain penalised a token **twice**: once on the raw
logits during the selection, then again after the temperature scaling. No llama.cpp chain does that —
the default chain lists `PENALTIES` exactly once, first, before the filters (the chain is built by
iterating `params.samplers`; reading the switch cases of `common_sampler_init` in their textual order
is what put the second pass after `TEMPERATURE`). In #53's two-token example the double pass shrank
the penalised token's probability from 10.50% to 2.55%.
## Commit 1 — the fix
- the final stage scales by `1/T` and nothing else: `scaled()` returns `sel_logit[i] * inv_t`
- the host reference runs the same single pass (it used to agree with the kernel's two-pass bug
instead of with llama.cpp), and fixture 9 pins the chain against the issue's independently derived
numbers — `0.10500059` for one pass vs the old `0.02550967` — plus a draw counter whose uniform
lands in the band where only the single-pass chain picks the penalised token
- docs and the header describe the single-pass order; setup requires engine 0.1.17 (0.1.16 is taken
by PR #54's engine)
## Commit 2 — a latent contract bug found while re-auditing the sampler
`use_bits` required only a non-null history, but the launch sizes the shared bitmap for
`penalty_last_n > 0` alone — a caller handing over a stale history buffer with the penalties
disabled made the kernel zero words of a **zero-byte** dynamic shared allocation. Unreachable from
the serve path (`set_history` passes `nullptr` when the window is empty) and invisible to memcheck,
but undefined per the contract of the public `sample_tokens`. The gate now also requires a
non-empty window (`hlen > 0`, both kernels); fixture 10 pins a stale history with `last_n = 0` to
the no-history result.
## Verification (RTX 4070 Ti SUPER, CUDA 13.4, real IQ3_XXS model)
- `sampler_parity --selftest`: 0 failures (fixtures 1–10, incl. #53's own numbers); each commit
built and tested individually
- randomised fuzz, kernel vs a serial single-pass reference: **2000 cases, 0 mismatches** (tie-heavy
logits, `top_k > n_vocab`, penalty windows longer/shorter than the buffer, `-1` padding,
seeds/counters above 2^32, batch rows)
- distribution test over 262,144 draws: empirical frequencies match `softmax(logits/T)` at
T ∈ {0.5, 0.7, 1.0, 2.0} within 5σ; `T = 0` is exactly the argmax; with penalties the empirical
distribution matches the single-pass softmax while the old two-pass one sits far outside tolerance
- `compute-sanitizer`: memcheck 0 errors, racecheck 0 hazards
- parity suite: 19/20 binaries pass (`ple_parity` needs generated PLE fixtures absent on this machine)
- end-to-end on the served IQ3_XXS model: `T=0` deterministic; seeded `T=0.6` reproduces exactly;
`T=1.4` differs; `T=0.05` shares a 220/220-char prefix with greedy; penalties change picks and cut
repeats (12→5); `min_p` works through the GEN line
- A/B vs the pre-fix engine on the real model: greedy and neutral-penalty sampling are bit-identical;
penalised sampling diverges exactly where it should站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。