Issues / #1736

#1736 Large context window presents a problem

open · @galmok · 1 comentários · No GitHub

Models & quants

Descrição

I am using opencode to work with Strata and while this seems fairly solid, I am hitting one problem again and again:

> prompt (230219 tokens) + max tokens (32000) exceeds the context (262144); requests are never truncated. Send a smaller max_tokens (at most 31917 here), or add "fit_max_tokens": true to the model's strata-<model>.json to shorten it to the room left (#545)

While this does show the problem and also presents a solution, I am uncertain what this does. Will it continue as normal (32000 token window) until the very end and then use a token window that fits the remaining space? Or will it change more than that?

If it just fits the last window to the remaining space, shouldn't that option be enabled by default?

No site

Links install, modelos, releases.