Pull requests / #1274
#1274 setup: --thinking / --instruct set the sampling every client gets (#1129)
open · @victorgabr · 0 コメント · GitHub で見る
本文
--thinking` / `--instruct` on setup (or at the start of an installed model, like --vram-reserve-mib) writes one of the two sampling sets the Qwen3.8-Flash-Next card recommends into the model's strata-<model>.json "sampling" block. A client that sends no sampling parameters of its own now gets those numbers instead of silently decoding greedy. - the point: traceability - the numbers in use are written down and named, instead of an unnamed "default" the way it was - fewer doom loops blamed on the engine: a request that sent no parameters decoded greedy (temperature=0) - the most repetition-prone setting - and nothing said so; the thinking default removes that, and the numbers are visible when a loop still happens - thinking is the default for new installs; hand-written numbers are kept unless a flag names otherwise - a start that replaces the block keeps the file as strata-<model>.json.bak; a no-op start writes nothing - setup's Settings line and the server's start line print the numbers and name the preset; a config without the block says requests decode greedily - greedy escape hatch: sampling.temperature = 0 - a "sampling" that is not a set of numbers (a name, a list, a bare value) stops the start with a message naming it; a setup run that replaces such a block says what it replaced Tests: 289 green across the setup and serve suites (incl. serve.test_server, 204).
関連リンク
インストール・モデル・リリースへの站内リンク。