Issues / #916
#916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
open · @lollo78 · 9 コメント · GitHub で見る
Setup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsSecurityDocumentationWindows
本文
**Hi, I'm trying to launch Strata also on my laptop, but it fails.
Here the full log:**
`Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
=== Step 1: checking your PC ===
[ok] GPU: NVIDIA GeForce RTX 2060, 6.0 GB VRAM, compute capability 7.5, driver 617.14
[!] less than 12 GB of VRAM: Strata will run, but most experts stay on the CPU and it will be slow
[ok] RAM: 32 GB
[ok] CPU: Intel(R) Core(TM) i7-10750H CPU @ 2.60GHz (AVX2)
=== Step 2: your choices ===
1) Qwen3.8-Flash-Next Qwen; GSQ-RCO quants by ISTA-DASLab - the original model
2) Swift 1.5 UkisAI's fine-tune of Qwen3.8-Flash-Next - thinks much shorter (-63% thinking tokens, 1.8x sooner answers by its authors' numbers)
3) Qwen3.8-Flash-Next Coder ISTA-DASLab's coding version - half the experts (code, tools, images kept): needs ~32 GB of RAM, faster; weaker outside coding
4) Qwen3.8-Flash-Next (Unsloth) Unsloth's ~4-bit quantizations - UD-IQ4_XS: a 94 GB download; with less than ~80 GB of RAM part of its experts are read from the SSD (UD-Q4_K_XL, 111 GB: experimental)
Which model? [1]: 3
[ok] model: Qwen3.8-Flash-Next Coder
1) IQ1_M the Coder's only size: half the experts, stored like IQ3_S (3.5 bits); download 58 GB, uses ~23 GB of RAM <- fits in the low-RAM mode (the GPU holds ~4%, the rest from the SSD)
Which size? [1]:
[ok] size: IQ1_M
Context length = how much text the model can see at once (your chat, files, tool output).
Longer needs more VRAM for it, so fewer experts fit on the GPU:
1) 8K tokens
2) 32K tokens (recommended for your GPU)
3) 64K tokens
4) 128K tokens
5) 256K tokens
6) 384K tokens (experimental: setup adds rope scaling)
7) 512K tokens (experimental: setup adds rope scaling)
Context? [2]: 2
[ok] context: 32768 tokens
KV cache precision (the model's memory of the conversation):
1) 8-bit (recommended: what every published number was measured with)
2) 4-bit half the memory (about 4% faster at 128K), but measurably less precise on long
documents; long-context lookups (needle tests) still pass
KV cache? [1]: 2
[ok] KV cache: 4-bit (Hadamard-rotated)
Images: the model can also read pictures (screenshots, photos, scanned pages). This adds a 0.9 GB
download and keeps ~1.4 GB of VRAM free for the image encoder, so text is a few % slower.
Do you want images? [n]: n
[ok] images: off
[ok] low-RAM mode: IQ1_M's experts (23 GB) are read from the model folder through the OS file cache instead of a copy in RAM (32 GB); the GPU holds ~4% of them
[!] most of the experts are read from the SSD while it answers: expect it to be much slower than with enough RAM (a faster SSD and a smaller size help)
EXPERIMENTAL - speed projection: a small control vector applied while the model runs (layers 4-44).
It changes how the model answers: its package describes it as a refusal-direction projection (the
model declines far fewer requests). Off unless you choose it; when on, the web app can switch it off
per chat. Details: data/experimental-speed-projection/README.md
Turn on the experimental speed projection? [n]:
[ok] experimental speed projection: off
=== Step 3: Python packages ===
Installing numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil ...
> D:\Strata-main\.venv\Scripts\python.exe -m pip install --quiet --disable-pip-version-check numpy==2.5.3; python_version >= "3.12" numpy==2.4.6; python_version == "3.11" numpy==2.2.6; python_version < "3.11" jinja2==3.1.6 regex==2026.9.10 pyyaml==6.0.3 tqdm==4.70.1 requests==2.34.2 cmake==4.4.3 ninja==1.13.2 pillow==12.3.0 psutil==7.2.2 markupsafe==3.0.3 certifi==2026.7.22 charset-normalizer==3.5.1 idna==3.20 urllib3==2.8.0 colorama==0.4.6; sys_platform == "win32"
[ok] numpy, jinja2, regex, pyyaml, tqdm, requests, cmake, ninja, pillow, psutil installed
=== Step 4: the Strata engine ===
llama.cpp source: 0.04 / 0.04 GB (100%)
[ok] llama.cpp source downloaded
[ok] llama.cpp 3cf0325 (gguf-py, ggml, mtmd)
Downloading the ready-made Strata engine ...
Strata engine: 0.09 / 0.12 GB (75%)
[ok] Strata engine downloaded
[ok] ready-made engine 0.1.39 for sm_75, sm_86, sm_89, sm_120 (CUDA 13.0)
Installing NVIDIA CUDA libraries (cuBLAS, CUDA runtime; ~0.4 GB) ...
> D:\Strata-main\.venv\Scripts\python.exe -m pip install --quiet --disable-pip-version-check nvidia-cublas==13.0.2.14 nvidia-cuda-runtime==13.0.96
[ok] NVIDIA CUDA libraries (cuBLAS, CUDA runtime; ~0.4 GB) installed
[ok] engine: D:\Strata-main\engine\strata.exe
=== Step 5: downloading Qwen3.8-Flash-Next Coder IQ1_M ===
The model files go in D:\Strata-data\models\coder-IQ1_M
Files you already have: put them here with their original names (Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf), or use --gguf-dir <their folder>.
Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf: 29.59 / 29.61 GB (100%)
[ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf downloaded
Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf: 28.78 / 28.80 GB (100%)
[ok] Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf downloaded
[ok] model files present
=== Step 6: preparing the model for Strata ===
> D:\Strata-main\.venv\Scripts\python.exe D:\Strata-main\tools\iq_pack.py --gguf D:\Strata-data\models\coder-IQ1_M\Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf --out D:\Strata-data\packs\coder-iq1_m
model shards: Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf
index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.37 GiB
tokenizer/: vocab 248320, merges 247587, pre qwen35, specials {'tokenizer.ggml.eos_token_id': 248046, 'tokenizer.ggml.padding_token_id': 248044, 'tokenizer.ggml.bos_token_id': 248044}
Writing the experts into one file for the low-RAM mode (one time, 23 GB) ...
> D:\Strata-main\.venv\Scripts\python.exe D:\Strata-main\tools\iq_pack.py --gguf D:\Strata-data\models\coder-IQ1_M\Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf --out D:\Strata-data\packs\coder-iq1_m --experts-bin
model shards: Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf, Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf
index.txt: 1079 tensors, 302 served natively, 0 in extra.bin, arena 1.37 GiB
layer 0 IQ3_XXS /IQ4_NL blob 2176000 at 0.00 GiB
layer 8 IQ2_S /Q2_0 blob 1510400 at 3.75 GiB
layer 16 IQ3_XXS /IQ4_NL blob 2176000 at 7.17 GiB
layer 24 IQ2_S /IQ4_NL blob 1971200 at 11.07 GiB
layer 32 IQ2_S /IQ4_NL blob 1971200 at 15.16 GiB
layer 40 IQ3_S /IQ4_NL blob 2329600 at 19.04 GiB
experts.bin: 48 layers, 23.42 GiB
[ok] model prepared: D:\Strata-data\packs\coder-iq1_m
The MTP draft layer (speculative decoding, ~2x faster output) comes from the original Qwen checkpoint:
only its ~5 GB of MTP tensors are downloaded.
> D:\Strata-main\.venv\Scripts\python.exe D:\Strata-main\tools\mtp_fetch.py fetch --out D:\Strata-data\mtp
# MTP block in the BF16 checkpoint
31 tensors, 5.214 GB, in 28 shards.
| tensor | dtype | shape | MB |
|---|---|---|---:|
| `mtp.fc_embedding.weight` | BF16 | 2560x2560 | 13.1 |
| `mtp.fc_hidden.weight` | BF16 | 2560x2560 | 13.1 |
| `mtp.hyper_connection_mixer.hc_norm.weight` | BF16 | 10240 | 0.0 |
| `mtp.hyper_connection_mixer.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
| `mtp.hyper_connection_mixer.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
| `mtp.layers.0.attn_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
| `mtp.layers.0.attn_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
| `mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
| `mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
| `mtp.layers.0.mlp.experts.down_proj` | BF16 | 512x2560x640 | 1677.7 |
| `mtp.layers.0.mlp.experts.gate_up_proj` | BF16 | 512x1280x2560 | 3355.4 |
| `mtp.layers.0.mlp.gate.weight` | BF16 | 512x2560 | 2.6 |
| `mtp.layers.0.mlp.shared_expert.down_proj.weight` | BF16 | 2560x640 | 3.3 |
| `mtp.layers.0.mlp.shared_expert.gate_proj.weight` | BF16 | 640x2560 | 3.3 |
| `mtp.layers.0.mlp.shared_expert.up_proj.weight` | BF16 | 640x2560 | 3.3 |
| `mtp.layers.0.mlp.shared_expert_gate.weight` | BF16 | 1x2560 | 0.0 |
| `mtp.layers.0.mlp_hyper_connection.block_inject_weight.weight` | BF16 | 4x10240 | 0.1 |
| `mtp.layers.0.mlp_hyper_connection.hc_norm.weight` | BF16 | 10240 | 0.0 |
| `mtp.layers.0.mlp_hyper_connection.input_mix_weight_down.weight` | BF16 | 320x10240 | 6.6 |
| `mtp.layers.0.mlp_hyper_connection.input_mix_weight_up.weight` | BF16 | 10240x320 | 6.6 |
| `mtp.layers.0.self_attn.indexer.index_qk_proj.weight` | BF16 | 640x2560 | 3.3 |
| `mtp.layers.0.self_attn.indexer.k_layernorm.weight` | BF16 | 128 | 0.0 |
| `mtp.layers.0.self_attn.indexer.q_layernorm.weight` | BF16 | 128 | 0.0 |
| `mtp.layers.0.self_attn.k_norm.weight` | BF16 | 256 | 0.0 |
| `mtp.layers.0.self_attn.k_proj.weight` | BF16 | 512x2560 | 2.6 |
| `mtp.layers.0.self_attn.o_proj.weight` | BF16 | 2560x6144 | 31.5 |
| `mtp.layers.0.self_attn.q_norm.weight` | BF16 | 256 | 0.0 |
| `mtp.layers.0.self_attn.q_proj.weight` | BF16 | 12288x2560 | 62.9 |
| `mtp.layers.0.self_attn.v_proj.weight` | BF16 | 512x2560 | 2.6 |
| `mtp.pre_fc_norm_embedding.weight` | BF16 | 2560 | 0.0 |
| `mtp.pre_fc_norm_hidden.weight` | BF16 | 10240 | 0.0 |
mtp.fc_embedding.weight 100%
mtp.fc_hidden.weight 100%
mtp.hyper_connection_mixer.hc_norm.weight 100%
mtp.hyper_connection_mixer.input_mix_weight_down.weight 100%
mtp.hyper_connection_mixer.input_mix_weight_up.weight 100%
mtp.layers.0.attn_hyper_connection.block_inject_weight.weight 100%
mtp.layers.0.attn_hyper_connection.hc_norm.weight 100%
mtp.layers.0.attn_hyper_connection.input_mix_weight_down.weight 100%
mtp.layers.0.attn_hyper_connection.input_mix_weight_up.weight 100%
mtp.layers.0.mlp.experts.down_proj 4%
mtp.layers.0.mlp.experts.down_proj 8%
mtp.layers.0.mlp.experts.down_proj 12%
mtp.layers.0.mlp.experts.down_proj 16%
mtp.layers.0.mlp.experts.down_proj 20%
mtp.layers.0.mlp.experts.down_proj 24%
mtp.layers.0.mlp.experts.down_proj 28%
mtp.layers.0.mlp.experts.down_proj 32%
mtp.layers.0.mlp.experts.down_proj 36%
mtp.layers.0.mlp.experts.down_proj 40%
mtp.layers.0.mlp.experts.down_proj 44%
mtp.layers.0.mlp.experts.down_proj 48%
mtp.layers.0.mlp.experts.down_proj 52%
mtp.layers.0.mlp.experts.down_proj 56%
mtp.layers.0.mlp.experts.down_proj 60%
mtp.layers.0.mlp.experts.down_proj 64%
mtp.layers.0.mlp.experts.down_proj 68%
mtp.layers.0.mlp.experts.down_proj 72%
mtp.layers.0.mlp.experts.down_proj 76%
mtp.layers.0.mlp.experts.down_proj 80%
mtp.layers.0.mlp.experts.down_proj 84%
mtp.layers.0.mlp.experts.down_proj 88%
mtp.layers.0.mlp.experts.down_proj 92%
mtp.layers.0.mlp.experts.down_proj 96%
mtp.layers.0.mlp.experts.down_proj 100%
mtp.layers.0.mlp.experts.gate_up_proj 2%
mtp.layers.0.mlp.experts.gate_up_proj 4%
mtp.layers.0.mlp.experts.gate_up_proj 6%
mtp.layers.0.mlp.experts.gate_up_proj 8%
mtp.layers.0.mlp.experts.gate_up_proj 10%
mtp.layers.0.mlp.experts.gate_up_proj 12%
mtp.layers.0.mlp.experts.gate_up_proj 14%
mtp.layers.0.mlp.experts.gate_up_proj 16%
mtp.layers.0.mlp.experts.gate_up_proj 18%
mtp.layers.0.mlp.experts.gate_up_proj 20%
mtp.layers.0.mlp.experts.gate_up_proj 22%
mtp.layers.0.mlp.experts.gate_up_proj 24%
mtp.layers.0.mlp.experts.gate_up_proj 26%
mtp.layers.0.mlp.experts.gate_up_proj 28%
mtp.layers.0.mlp.experts.gate_up_proj 30%
mtp.layers.0.mlp.experts.gate_up_proj 32%
mtp.layers.0.mlp.experts.gate_up_proj 34%
mtp.layers.0.mlp.experts.gate_up_proj 36%
mtp.layers.0.mlp.experts.gate_up_proj 38%
mtp.layers.0.mlp.experts.gate_up_proj 40%
mtp.layers.0.mlp.experts.gate_up_proj 42%
mtp.layers.0.mlp.experts.gate_up_proj 44%
mtp.layers.0.mlp.experts.gate_up_proj 46%
mtp.layers.0.mlp.experts.gate_up_proj 48%
mtp.layers.0.mlp.experts.gate_up_proj 5関連リンク
インストール・モデル・リリースへの站内リンク。