Issues / #841
#841 not stable start with 3+ cards
open · @yotadrivers-blip · 8 comments · View on GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Description
hello! about 10-15 starts and only 2 sucsees. Getting error out of memory for 1 2 or all cards.
try: fixex expert cashe, placing cache at 3rd card, spliting layers like 16,32.
able to start only with 2 cards option w/o vision and not stable.
3060-12gb X3
json file auto created with 2 cars choose have strata-swift-iq2_xs.json "gpu": [ 0, 1, 2 ]
re running setup causing sometimes auto filling options like 1 1 2 2 1 e.t.c without keys press
[strata] experts loaded: 33.02 GiB at 0.07 GiB/s (579 s so far)
[strata] filling the GPU's expert cache (4884 experts, 6.33 GiB of VRAM) ...
[strata] still starting (599 s) - please wait ...
Traceback (most recent call last):
File "F:\!strada\serve\server.py", line 4145, in <module>
sys.exit(main())
~~~~^^
File "F:\!strada\serve\server.py", line 4001, in main
engine = StrataEngine(exe, engine_args(cfg) + (effort_end or []), cwd=cfg.get("cwd"), log=cfg.get("log"),
env=env, lazy=lazy)
File "F:\!strada\serve\server.py", line 478, in __init__
raise RuntimeError("the engine exited before it was ready" + (f" (see {log})" if log else "") +
start_failure_hint(log, log_start) + start_log_tail(log, log_start))
RuntimeError: the engine exited before it was ready (see F:\!strada\strata-swift-iq2_xs.log)
the engine log's last lines:
strata generate: PLE on, table 320001536 rows of F:\!strada\Strata-data\models\swift-IQ2_XS\Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf
strata generate: layer split: CUDA1 holds its weights, session [16, 33); 7.79 GiB free
strata generate: layer split: CUDA2 holds its weights, session [33, 48) and the head; 7.47 GiB free
strata mtp: draft layer loaded, 839 MiB of VRAM (experts 675, dense 111), files read in 9.59 s (82 MiB/s)
strata generate: GPU 0: NVIDIA GeForce RTX 3060, compute capability 8.6
strata generate: multi-GPU under WDDM: at most 8 GiB of the expert arena is pinned (STRATA_ARENA_PIN_GIB changes it)
strata generate: expert arena read unbuffered (0 of 32 probe reads from the file cache; 16.8 GiB available, 63.5 GiB of files)
strata generate: expert arena: locked 25626 MiB via working-set minimum + VirtualLock; cudaHostRegister limited to 8 GiB by the engine (multi-GPU under WDDM, or remote experts); 12 slices pinned (7 GiB); large pages refused for 35456548864 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata generate: loaded 33.02 GiB at 0.07 GiB/s
strata generate: hint: ~24x below what this hardware streams from a normal launch. If Strata is started by Task Scheduler or a service, register the task with Priority 4 (Normal) and 'Run with highest privileges' - the scheduler's defaults (Below normal + a least-privilege token) throttle the load. See docs/DETAILS.md ('Running it at startup').
strata generate: expert cache 4884 slots, 6.33 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
so a reply can differ slightly from a run without the cache (same quality:
bench/results/2026-09-27-cache-parity).
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 4884 of 4884 slots from the profile; slot 0 verified
strata generate: layer split, CUDA1: 7.79 GiB free of 12.00, room for experts 7.01 GiB
strata generate: layer split: CUDA1 runs layers 16-32, expert cache 5198 slots (7.01 GiB), 5198 of its 8704 profiled pairs; slot 0 verified
strata generate: layer split, CUDA2: 6.65 GiB free of 12.00, room for experts 5.87 GiB
strata generate: layer split, CUDA2 expert cache: ExpertCache: cudaMalloc(5.87 GiB) failed: out of memoryRelated on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.