Issues / #1550

#1550 A question/issue on ToolCall

open · @tamascogustavo · 0 Kommentare · Auf GitHub

Setup & installNVIDIA / CUDAModels & quants

Beschreibung

First of all thanks for such an amazing project ! This has improved a lot my local usage of the model.

I am running strata with the following optimized parameters:
```
./setup.sh
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
  [ok] GPU: GPU 0 (NVIDIA GeForce RTX 5070, 12 GB)

  ----------------------------------------------------------------------------------------------------
  Starting qwen3.8-flash-next-iq3_s: it loads about 55 GB into RAM and locks part of it for the GPU.
  While it does, YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES (longer the first time after a
  restart). That is normal: please wait and don't close this window - the browser opens when it is ready.
  Later, closing this window stops the model.
  ----------------------------------------------------------------------------------------------------
  Settings (strata-iq3_s.json): --expert-cache auto --prefill auto --spec 4 --max-context 131072 --kv
    int8 --pcie-frac 0.35 --spec-min-p 0.70; server 127.0.0.1:8080, gpu 0
loading the model (the first start takes a minute or two) ...
```
And i notice that many times when the model is loaded in the engine, the tool call do not work and the model behaves really bad, with more prominent loops and the quality of the reasoning is also affected.

Upon stop and restarting the engine a few times, all works fine. But no warnings  or erros during loading appear in terminal logs. 

Mehr auf der Site

Links zu Install, Modellen, Releases.