Issues / #984

#984 reasoning_budget_tokens is ignored?

open · @PentaForgeDev · 5 comments · View on GitHub

Setup & installModels & quants

Description

Heyo.  Just downloaded, so v0.1.39.  Ran the setup.sh on uBuntu 26, all good.

[strata] logging output shows 'thinking: x of max 6400 tokens'.  I set `"reasoning_budget_tokens": 2048` in the config, but it still thinks up to 6400 tokens, rather than 2048 as set in the config:

```
{
 "exe": "...",
 "args": [
  "--pack",
  "~/devl/AI/Strata-data/packs/iq2_xs",
  "--native",
  "~/devl/AI/Strata-data/models/IQ2_XS/Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf",
  "--ple-gguf",
  "~/devl/AI/Strata-data/models/IQ2_XS/Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00002-of-00002.gguf",
  "--expert-profile",
  "~/devl/AI/Strata/data/expert-profile.bin",
  "--expert-cache",
  "auto",
  "--prefill",
  "auto",
  "--spec",
  "4",
  "--mtp",
  "~/devl/AI/Strata-data/mtp/rt",
  "--max-context",
  "65535",
  "--kv",
  "int8",
  "--spec-min-p",
  "0.5",
  "--pool-workers",
  "10"
],
 "model_name": "qwen3.8-flash-next-iq2_xs",
 "port": 11433,
 "gpu": 0,
 "gpus_asked": true,
 "reasoning_budget_tokens": 2048
}
```

Interestingly, the logging also shows `answering: x of max 6400 tokens`, so im not sure where the 6400 is coming from?

Right now, the model just fills itself up thinking then fails to answer.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.