Issues / #1445

#1445 vision encoder did not start on V100-SXM2-32GB

closed · @TAIIOK · 2 comments · View on GitHub

Setup & installServer & APINVIDIA / CUDAModels & quantsLinux

Description

Vision model fails to start on Tesla V100-SXM2-32GB
Environment
GPU: NVIDIA Tesla V100-SXM2 32 GB
OS: Linux Ubuntu 24.04
Model:  swift-IQ3_XXS
CUDA: 12.9  
Vision: enabled
Context: 131072
KV cache: int8
Problem

The vision encoder fails to start.

I get the following error when running calibration:

Command
./setup.sh --calibrate
Output
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)

  Tuning Strata for this PC: the output speed is measured with a few engine settings (the PCIe share, the
  draft depth, the CPU threads, the expert cache). It takes about 10 minutes; the PC is busy meanwhile.
  Loading the model for the measurements ...
[strata] starting the engine: reading the model's weights ...
  [!]  the tuning did not finish (the engine exited before it was ready (see /home/grapeuser/Strata/strata-swift-iq3_xxs.log)
the engine log's last lines:
  [strata] 2026-10-07 20:46:45 engine started: /home/grapeuser/Strata/engine-cuda12/strata --serve --pack /home/grapeuser/Strata-data/packs/swift-iq3_xxs --native /home/grapeuser/Strata-data/models/swift-IQ3_XXS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf --ple-gguf /home/grapeuser/Strata-data/models/swift-IQ3_XXS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf --expert-profile /home/grapeuser/Strata/data/expert-profile.bin --expert-cache auto --prefill auto --spec 4 --mtp /home/grapeuser/Strata-data/mtp/rt --max-context 131072 --kv int8 --vision --vram-reserve-mib 700 --spec-min-p 0.5
  strata generate: PCIe probe failed -> pcie_frac default 0.55
  strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
  strata generate: native embedding: cannot pin 260 MiB: CUDA-capable device(s) is/are busy or unavailable, nor place it in VRAM): the default settings stay
       the engine said: strata generate: native embedding: cannot pin 260 MiB: CUDA-capable device(s) is/are busy or unavailable, nor place it in VRAM

  [!]  this PC is NOT tuned: the tuning failed (the reason is above); the model starts with the default settings
  [ok] GPU: GPU 0 (Tesla V100-SXM2-32GB, 32 GB)

  ----------------------------------------------------------------------------------------------------
  Starting swift-1.5-iq3_xxs: it loads about 40 GB into RAM and locks part of it for the GPU.
  While it does, YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES (longer the first time after a
  restart). That is normal: please wait and don't close this window - the browser opens when it is ready.
  Later, closing this window stops the model.
  ----------------------------------------------------------------------------------------------------

  Settings (strata-swift-iq3_xxs.json): --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5
    --max-context 131072 --kv int8 --vision --vram-reserve-mib 700; server 127.0.0.1:8080, gpu 0

loading the vision encoder ...

Traceback (most recent call last):
  File "/home/grapeuser/Strata/serve/server.py", line 5611, in <module>
    sys.exit(main())
             ^^^^^^
  File "/home/grapeuser/Strata/serve/server.py", line 5411, in main
    vision = Vision(vcfg, log=open(cfg["log"], "a", encoding="utf-8") if cfg.get("log") else None,
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/grapeuser/Strata/serve/server.py", line 1741, in __init__
    self._start()
  File "/home/grapeuser/Strata/serve/server.py", line 1763, in _start
    raise RuntimeError("the vision encoder did not start: " + line.strip())
RuntimeError: the vision encoder did not start:

ggml_backend_cuda_device_get_memory: cudaMemGetInfo failed (CUDA-capable device(s) is/are busy or unavailable), returning 0/0
ggml_backend_cuda_device_get_memory: cudaMemGetInfo failed (CUDA-capable device(s) is/are busy or unavailable), returning 0/0
load_hparams: Qwen-VL models require at minimum 1024 image tokens to function correctly on grounding tasks
load_hparams: if you encounter problems with accuracy, try adding --image-min-tokens 1024
load_hparams: more info: https://github.com/ggml-org/llama.cpp/issues/16842


In log file: strata-swift-iq3_xxs.log  
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 865.48 MiB on device 0: cudaMalloc failed: CUDA-capable device(s) is/are busy or unavailable
alloc_tensor_range: failed to allocate CUDA0 buffer of size 907524736
/home/grapeuser/Strata/third_party/llama.cpp/ggml/src/ggml-backend.cpp:188: GGML_ASSERT(buffer) failed
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x136e70b)[0x5be57a06170b]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x136ecfc)[0x5be57a061cfc]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x136eed7)[0x5be57a061ed7]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x138cbec)[0x5be57a07fbec]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x1bd178)[0x5be578eb0178]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x1a4bb5)[0x5be578e97bb5]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x103a97)[0x5be578df6a97]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0xfb2b1)[0x5be578dee2b1]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0x6e1fa)[0x5be578d611fa]
/lib/x86_64-linux-gnu/libc.so.6(+0x2a1ca)[0x7f15b702a1ca]
/lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0x8b)[0x7f15b702a28b]
/home/grapeuser/Strata/engine-cuda12/strata-vision(+0xf7485)[0x5be578dea485]

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.