Issues / #694

#694 fit_max_tokens issue, context problem with agentic coding.

closed · @mkultra333 · 3 comentários · No GitHub

Setup & installModels & quants

Descrição

I'm using Strata with Oh My Pi, context size 131072 and maxTokens in OMP set to 64000.  I would run into a problem where the existing context plus the next promt size would choke.  The error message from OMP is: 
```

✘ 400 prompt (67811 tokens) + max tokens (64000) exceeds the context (131072); requests are never
 truncated. Se…
   prompt (67811 tokens) + max tokens (64000) exceeds the context (131072); requests are never
 truncated. Send a…
   raw-http-request=C:\Users\Jared\.omp\logs\http-400-requests\1791019190383-1s8auzy6vao75.json
 Dismissed when you send your next message.

```
After some back and forth with Claude we found that this can be fixed with a setting already in Strata, change fit_max_tokens from False to True. I can't find any way to do this manually though, if I add it to my strata-iq3_xxs.json file it just gets overwritten again when I launch.

The only way I found to change it was in setup.py, find this code:
```

    cfg = {"exe": str(eng / EXE), "args": args, "cwd": str(ROOT), "tokenizer": str(pack / "tokenizer"),
           "model_name": f"{fam['name']}-{model.lower()}", "log": str(ROOT / f"strata-{tag.lower()}.log"),
           "lib_dirs": lib_dirs, "port": a.port}

```
and change the last line to:

```
           "lib_dirs": lib_dirs, "port": a.port, "fit_max_tokens": True}
```

After that, 
"fit_max_tokens": true,
appears in strata-iq3_xxs.json and the context problem in OMP goes away and it all works fine.

Perhaps fit_max_tokens should be true by default?  Or easily set with a command line option?

No site

Links install, modelos, releases.