Issues / #1550

#1550 A question/issue on ToolCall

open · @tamascogustavo · 0 comments · View on GitHub

Setup & installNVIDIA / CUDAModels & quants

Description

First of all thanks for such an amazing project ! This has improved a lot my local usage of the model.

I am running strata with the following optimized parameters:
```
./setup.sh
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
  [ok] GPU: GPU 0 (NVIDIA GeForce RTX 5070, 12 GB)

  ----------------------------------------------------------------------------------------------------
  Starting qwen3.8-flash-next-iq3_s: it loads about 55 GB into RAM and locks part of it for the GPU.
  While it does, YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES (longer the first time after a
  restart). That is normal: please wait and don't close this window - the browser opens when it is ready.
  Later, closing this window stops the model.
  ----------------------------------------------------------------------------------------------------
  Settings (strata-iq3_s.json): --expert-cache auto --prefill auto --spec 4 --max-context 131072 --kv
    int8 --pcie-frac 0.35 --spec-min-p 0.70; server 127.0.0.1:8080, gpu 0
loading the model (the first start takes a minute or two) ...
```
And i notice that many times when the model is loaded in the engine, the tool call do not work and the model behaves really bad, with more prominent loops and the quality of the reasoning is also affected.

Upon stop and restarting the engine a few times, all works fine. But no warnings  or erros during loading appear in terminal logs. 

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.