Issues / #916

#916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM

open · @lollo78 · 9 评论 · 在 GitHub 查看

Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsSecurityDocumentationWindows

描述

**Hi, I'm trying to launch Strata also on my laptop, but it fails.

Here the full log:**

`Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)

=== Step 1: checking your PC ===
  [ok] GPU: NVIDIA GeForce RTX 2060, 6.0 GB VRAM, compute capability 7.5, driver 617.14
  [!]  less than 12 GB of VRAM: Strata will run, but most experts stay on the CPU and it will be slow
  [ok] RAM: 32 GB
  [ok] CPU: Intel(R) Core(TM) i7-10750H CPU @ 2.60GHz (AVX2)

=== Step 2: your choices ===
  1) Qwen3.8-Flash-Next   Qwen; GSQ-RCO quants by ISTA-DASLab - the original model
  2) Swift 1.5            UkisAI's fine-tune of Qwen3.8-Flash-Next - thinks much shorter (-63% thinking tokens, 1.8x sooner answers by its authors' numbers)
  3) Qwen3.8-Flash-Next Coder ISTA-DASLab's coding version - half the experts (code, tools, images kept): needs ~32 GB of RAM, faster; weaker outside coding
  4) Qwen3.8-Flash-Next (Unsloth) Unsloth's ~4-bit quantizations - UD-IQ4_XS: a 94 GB download; with less than ~80 GB of RAM part of its experts are read from the SSD (UD-Q4_K_XL, 111 GB: experimental)
Which model? [1]: 3
  [ok] model: Qwen3.8-Flash-Next Coder

  1) IQ1_M    the Coder's only size: half the experts, stored like IQ3_S (3.5 bits); download 58 GB, uses ~23 GB of RAM   <- fits in the low-RAM mode (the GPU holds ~4%, the rest from the SSD)
Which size? [1]:
  [ok] size: IQ1_M

  Context length = how much text the model can see at once (your chat, files, tool output).
  Longer needs more VRAM for it, so fewer experts fit on the GPU:
  1) 8K tokens
  2) 32K tokens   (recommended for your GPU)
  3) 64K tokens
  4) 128K tokens
  5) 256K tokens
  6) 384K tokens   (experimental: setup adds rope scaling)
  7) 512K tokens   (experimental: setup adds rope scaling)
Context? [2]: 2
  [ok] context: 32768 tokens

  KV cache precision (the model's memory of the conversation):
  1) 8-bit   (recommended: what every published number was measured with)
  2) 4-bit   half the memory (about 4% faster at 128K), but measurably less precise on long
             documents; long-context lookups (needle tests) still pass
KV cache? [1]: 2
  [ok] KV cache: 4-bit (Hadamard-rotated)

  Images: the model can also read pictures (screenshots, photos, scanned pages). This adds a 0.9 GB
  download and keeps ~1.4 GB of VRAM free for the image encoder, so text is a few % slower.
Do you want images? [n]: n
  [ok] images: off
  [ok] low-RAM mode: IQ1_M's experts (23 GB) are read from the model folder through the OS file cache instead of a copy in RAM (32 GB); the GPU holds ~4% of them
  [!]  most of the experts are read from the SSD while it answers: expect it to be much slower than with enough RAM (a faster SSD and a smaller size help)

  EXPERIMENTAL - speed projection: a small control vector applied while the model runs (layers 4-44).
  It changes how the model answers: its package describes it as a refusal-direction projection (the
  model declines far fewer requests). Off unless you choose it; when on, the web app can switch it off
  per chat. Details: data/experimental-speed-projection/README.md
Turn on the experimental speed projection? [n]:
  [ok] experimental speed projection: off

=== Step 3: Python packages ===
  Installing numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil ...
  > D:\Strata-main\.venv\Scripts\python.exe -m pip install --quiet --disable-pip-version-check numpy==2.5.3; python_version >= "3.12" numpy==2.4.6; python_version == "3.11" numpy==2.2.6; python_version < "3.11" jinja2==3.1.6 regex==2026.9.10 pyyaml==6.0.3 tqdm==4.70.1 requests==2.34.2 cmake==4.4.3 ninja==1.13.2 pillow==12.3.0 psutil==7.2.2 markupsafe==3.0.3 certifi==2026.7.22 charset-normalizer==3.5.1 idna==3.20 urllib3==2.8.0 colorama==0.4.6; sys_platform == "win32"
  [ok] numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil installed

=== Step 4: the Strata engine ===
  llama.cpp source:   0.04 / 0.04 GB (100%)
  [ok] llama.cpp source downloaded
  [ok] llama.cpp 3cf0325 (gguf-py, ggml, mtmd)
  Downloading the ready-made Strata engine ...
  Strata engine:   0.09 / 0.12 GB (75%)
  [ok] Strata engine downloaded
  [ok] ready-made engine 0.1.39 for sm_75, sm_86, sm_89, sm_120 (CUDA 13.0)
  Installing NVIDIA CUDA libraries (cuBLAS, CUDA runtime; ~0.4 GB) ...
  > D:\Strata-main\.venv\Scripts\python.exe -m pip install --quiet --disable-pip-version-check nvidia-cublas==13.0.2.14 nvidia-cuda-runtime==13.0.96
  [ok] NVIDIA CUDA libraries (cuBLAS, CUDA runtime; ~0.4 GB) installed
  [ok] engine: D:\Strata-main\engine\strata.exe

=== Step 5: downloading Qwen3.8-Flash-Next Coder IQ1_M ===
  The model files go in D:\Strata-data\models\coder-IQ1_M
  Files you already have: put them here with their original names (Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf), or use --gguf-dir <their folder>.
  Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf:  29.59 / 29.61 GB (100%)
  [ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf downloaded
  Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf:  28.78 / 28.80 GB (100%)
  [ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf downloaded
  [ok] model files present

=== Step 6: preparing the model for Strata ===
  > D:\Strata-main\.venv\Scripts\python.exe D:\Strata-main\tools\iq_pack.py --gguf D:\Strata-data\models\coder-IQ1_M\Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf --out D:\Strata-data\packs\coder-iq1_m
model shards: Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf
index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.37 GiB
tokenizer/: vocab 248320, merges 247587, pre qwen35, specials {'tokenizer.ggml.eos_token_id': 248046, 'tokenizer.ggml.padding_token_id': 248044, 'tokenizer.ggml.bos_token_id': 248044}
  Writing the experts into one file for the low-RAM mode (one time, 23 GB) ...
  > D:\Strata-main\.venv\Scripts\python.exe D:\Strata-main\tools\iq_pack.py --gguf D:\Strata-data\models\coder-IQ1_M\Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf --out D:\Strata-data\packs\coder-iq1_m --experts-bin
model shards: Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf
index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.37 GiB
  layer  0  IQ3_XXS /IQ4_NL  blob  2176000  at 0.00 GiB
  layer  8  IQ2_S   /Q2_0    blob  1510400  at 3.75 GiB
  layer 16  IQ3_XXS /IQ4_NL  blob  2176000  at 7.17 GiB
  layer 24  IQ2_S   /IQ4_NL  blob  1971200  at 11.07 GiB
  layer 32  IQ2_S   /IQ4_NL  blob  1971200  at 15.16 GiB
  layer 40  IQ3_S   /IQ4_NL  blob  2329600  at 19.04 GiB
experts.bin: 48 layers, 23.42 GiB
  [ok] model prepared: D:\Strata-data\packs\coder-iq1_m
  The MTP draft layer (speculative decoding, ~2x faster output) comes from the original Qwen checkpoint:
  only its ~5 GB of MTP tensors are downloaded.
  > D:\Strata-main\.venv\Scripts\python.exe D:\Strata-main\tools\mtp_fetch.py fetch --out D:\Strata-data\mtp
# MTP block in the BF16 checkpoint

31 tensors, 5.214 GB, in 28 shards.

| tensor | dtype | shape | MB |
|---|---|---|---:|
| `mtp.fc_embedding.weight` | BF16 | 2560x2560 | 13.1 |
| `mtp.fc_hidden.weight` | BF16 | 2560x2560 | 13.1 |
| `mtp.hyper_connection_mixer.hc_norm.weight` | BF16 | 10240 | 0.0 |
| `mtp.hyper_connection_mixer.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
| `mtp.hyper_connection_mixer.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
| `mtp.layers.0.attn_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
| `mtp.layers.0.attn_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
| `mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
| `mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
| `mtp.layers.0.mlp.experts.down_proj` | BF16 | 512x2560x640 | 1677.7 |
| `mtp.layers.0.mlp.experts.gate_up_proj` | BF16 | 512x1280x2560 | 3355.4 |
| `mtp.layers.0.mlp.gate.weight` | BF16 | 512x2560 | 2.6 |
| `mtp.layers.0.mlp.shared_expert.down_proj.weight` | BF16 | 2560x640 | 3.3 |
| `mtp.layers.0.mlp.shared_expert.gate_proj.weight` | BF16 | 640x2560 | 3.3 |
| `mtp.layers.0.mlp.shared_expert.up_proj.weight` | BF16 | 640x2560 | 3.3 |
| `mtp.layers.0.mlp.shared_expert_gate.weight` | BF16 | 1x2560 | 0.0 |
| `mtp.layers.0.mlp_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
| `mtp.layers.0.mlp_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
| `mtp.layers.0.mlp_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
| `mtp.layers.0.mlp_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
| `mtp.layers.0.self_attn.indexer.index_qk_proj.weight` | BF16 | 640x2560 | 3.3 |
| `mtp.layers.0.self_attn.indexer.k_layernorm.weight` | BF16 | 128 | 0.0 |
| `mtp.layers.0.self_attn.indexer.q_layernorm.weight` | BF16 | 128 | 0.0 |
| `mtp.layers.0.self_attn.k_norm.weight` | BF16 | 256 | 0.0 |
| `mtp.layers.0.self_attn.k_proj.weight` | BF16 | 512x2560 | 2.6 |
| `mtp.layers.0.self_attn.o_proj.weight` | BF16 | 2560x6144 | 31.5 |
| `mtp.layers.0.self_attn.q_norm.weight` | BF16 | 256 | 0.0 |
| `mtp.layers.0.self_attn.q_proj.weight` | BF16 | 12288x2560 | 62.9 |
| `mtp.layers.0.self_attn.v_proj.weight` | BF16 | 512x2560 | 2.6 |
| `mtp.pre_fc_norm_embedding.weight` | BF16 | 2560 | 0.0 |
| `mtp.pre_fc_norm_hidden.weight` | BF16 | 10240 | 0.0 |

mtp.fc_embedding.weight 100%
mtp.fc_hidden.weight 100%
mtp.hyper_connection_mixer.hc_norm.weight 100%
mtp.hyper_connection_mixer.input_mix_weight_down.weight 100%
mtp.hyper_connection_mixer.input_mix_weight_up.weight 100%
mtp.layers.0.attn_hyper_connection.block_inject_weight.weight 100%
mtp.layers.0.attn_hyper_connection.hc_norm.weight 100%
mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight 100%
mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight 100%
mtp.layers.0.mlp.experts.down_proj 4%
mtp.layers.0.mlp.experts.down_proj 8%
mtp.layers.0.mlp.experts.down_proj 12%
mtp.layers.0.mlp.experts.down_proj 16%
mtp.layers.0.mlp.experts.down_proj 20%
mtp.layers.0.mlp.experts.down_proj 24%
mtp.layers.0.mlp.experts.down_proj 28%
mtp.layers.0.mlp.experts.down_proj 32%
mtp.layers.0.mlp.experts.down_proj 36%
mtp.layers.0.mlp.experts.down_proj 40%
mtp.layers.0.mlp.experts.down_proj 44%
mtp.layers.0.mlp.experts.down_proj 48%
mtp.layers.0.mlp.experts.down_proj 52%
mtp.layers.0.mlp.experts.down_proj 56%
mtp.layers.0.mlp.experts.down_proj 60%
mtp.layers.0.mlp.experts.down_proj 64%
mtp.layers.0.mlp.experts.down_proj 68%
mtp.layers.0.mlp.experts.down_proj 72%
mtp.layers.0.mlp.experts.down_proj 76%
mtp.layers.0.mlp.experts.down_proj 80%
mtp.layers.0.mlp.experts.down_proj 84%
mtp.layers.0.mlp.experts.down_proj 88%
mtp.layers.0.mlp.experts.down_proj 92%
mtp.layers.0.mlp.experts.down_proj 96%
mtp.layers.0.mlp.experts.down_proj 100%
mtp.layers.0.mlp.experts.gate_up_proj 2%
mtp.layers.0.mlp.experts.gate_up_proj 4%
mtp.layers.0.mlp.experts.gate_up_proj 6%
mtp.layers.0.mlp.experts.gate_up_proj 8%
mtp.layers.0.mlp.experts.gate_up_proj 10%
mtp.layers.0.mlp.experts.gate_up_proj 12%
mtp.layers.0.mlp.experts.gate_up_proj 14%
mtp.layers.0.mlp.experts.gate_up_proj 16%
mtp.layers.0.mlp.experts.gate_up_proj 18%
mtp.layers.0.mlp.experts.gate_up_proj 20%
mtp.layers.0.mlp.experts.gate_up_proj 22%
mtp.layers.0.mlp.experts.gate_up_proj 24%
mtp.layers.0.mlp.experts.gate_up_proj 26%
mtp.layers.0.mlp.experts.gate_up_proj 28%
mtp.layers.0.mlp.experts.gate_up_proj 30%
mtp.layers.0.mlp.experts.gate_up_proj 32%
mtp.layers.0.mlp.experts.gate_up_proj 34%
mtp.layers.0.mlp.experts.gate_up_proj 36%
mtp.layers.0.mlp.experts.gate_up_proj 38%
mtp.layers.0.mlp.experts.gate_up_proj 40%
mtp.layers.0.mlp.experts.gate_up_proj 42%
mtp.layers.0.mlp.experts.gate_up_proj 44%
mtp.layers.0.mlp.experts.gate_up_proj 46%
mtp.layers.0.mlp.experts.gate_up_proj 48%
mtp.layers.0.mlp.experts.gate_up_proj 5

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。