Issues / #530
#530 high/xhigh reasoning_effort can silently run to max_tokens with empty content when reasoning_budget_tokens isn't set
closed · @maxlippe-leonardo · 1 コメント · GitHub で見る
Setup & installMulti-GPUNVIDIA / CUDAWindows
本文
#123 added the fix, but nothing defaults it or warns about it, and it's easy to not know you need it. 5 separate runs here (same prompt, temp 0, seed 42) hit `finish_reason: length` with empty `content` — reasoning alone filled the whole budget, never got to the answer. Happened across two engine versions (0.1.32, 0.1.35), with and without a system message, with `xhigh` and the official `high`. Setting `reasoning_budget_tokens` fixed it in our repro (`finish_reason: stop`, valid content, actually fewer total tokens than any of the failed runs). Worth noting for honesty: one earlier run at the same effort level with a system message *did* finish normally without the parameter set, so this isn't 100% deterministic every time at temp 0 — but it was common enough across our testing to be a real footgun, not a one-off. Suggestion: either a sane default tied to `reasoning_effort` (especially `high`/`xhigh`), or at minimum a server-side warning/log line when a response finishes with `length` and empty content, pointing at `reasoning_budget_tokens`. Setup: Windows, dual GPU (RTX 3080 Ti 12GB + RTX 3080 10GB), Swift 1.5 IQ2_XS.
関連リンク
インストール・モデル・リリースへの站内リンク。