Issues / #1736

#1736 Large context window presents a problem

open · @galmok · 1 comments · View on GitHub

Models & quants

Description

I am using opencode to work with Strata and while this seems fairly solid, I am hitting one problem again and again:

> prompt (230219 tokens) + max tokens (32000) exceeds the context (262144); requests are never truncated. Send a smaller max_tokens (at most 31917 here), or add "fit_max_tokens": true to the model's strata-<model>.json to shorten it to the room left (#545)

While this does show the problem and also presents a solution, I am uncertain what this does. Will it continue as normal (32000 token window) until the very end and then use a token window that fits the remaining space? Or will it change more than that?

If it just fits the last window to the remaining space, shouldn't that option be enabled by default?

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.