Issues / #984

#984 reasoning_budget_tokens is ignored?

open · @PentaForgeDev · 5 评论 · 在 GitHub 查看

Setup & installModels & quants

描述

Heyo.  Just downloaded, so v0.1.39.  Ran the setup.sh on uBuntu 26, all good.

[strata] logging output shows 'thinking: x of max 6400 tokens'.  I set `"reasoning_budget_tokens": 2048` in the config, but it still thinks up to 6400 tokens, rather than 2048 as set in the config:

```
{
 "exe": "...",
 "args": [
  "--pack",
  "~/devl/AI/Strata-data/packs/iq2_xs",
  "--native",
  "~/devl/AI/Strata-data/models/IQ2_XS/Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf",
  "--ple-gguf",
  "~/devl/AI/Strata-data/models/IQ2_XS/Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00002-of-00002.gguf",
  "--expert-profile",
  "~/devl/AI/Strata/data/expert-profile.bin",
  "--expert-cache",
  "auto",
  "--prefill",
  "auto",
  "--spec",
  "4",
  "--mtp",
  "~/devl/AI/Strata-data/mtp/rt",
  "--max-context",
  "65535",
  "--kv",
  "int8",
  "--spec-min-p",
  "0.5",
  "--pool-workers",
  "10"
],
 "model_name": "qwen3.8-flash-next-iq2_xs",
 "port": 11433,
 "gpu": 0,
 "gpus_asked": true,
 "reasoning_budget_tokens": 2048
}
```

Interestingly, the logging also shows `answering: x of max 6400 tokens`, so im not sure where the 6400 is coming from?

Right now, the model just fills itself up thinking then fails to answer.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。