Pull requests / #1274

#1274 setup: --thinking / --instruct set the sampling every client gets (#1129)

open · @victorgabr · 0 commentaires · Sur GitHub

Setup & installServer & API

Description

--thinking` / `--instruct` on setup (or at the start of an installed model, like
--vram-reserve-mib) writes one of the two sampling sets the Qwen3.8-Flash-Next card
recommends into the model's strata-<model>.json "sampling" block. A client that sends
no sampling parameters of its own now gets those numbers instead of silently decoding
greedy.

- the point: traceability - the numbers in use are written down and named, instead of
  an unnamed "default" the way it was
- fewer doom loops blamed on the engine: a request that sent no parameters decoded
  greedy (temperature=0) - the most repetition-prone setting - and nothing said so;
  the thinking default removes that, and the numbers are visible when a loop still
  happens
- thinking is the default for new installs; hand-written numbers are kept unless a
  flag names otherwise
- a start that replaces the block keeps the file as strata-<model>.json.bak; a no-op
  start writes nothing
- setup's Settings line and the server's start line print the numbers and name the
  preset; a config without the block says requests decode greedily
- greedy escape hatch: sampling.temperature = 0
- a "sampling" that is not a set of numbers (a name, a list, a bare value) stops the
  start with a message naming it; a setup run that replaces such a block says what
  it replaced

Tests: 289 green across the setup and serve suites (incl. serve.test_server, 204).

Sur le site

Liens install, modèles, releases.