Issues / #1094

#1094 --layer-split auto fails on dual GPUs for 512k context

open · @fukc-gihtub · 1 commentaires · Sur GitHub

BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quants

Description

With `"layer-split": "auto"`, Strata v0.1.40.1 OOMs on my machine (see the log below). When changed to `"layer_split": "46"`, it loads and responds.

```
$ bash run-swift-iq3_xxs.sh 
loading the model (the first start takes a minute or two) ...
[strata] layer split across GPUs [0, 1] (auto)
[strata] starting the engine: reading the model's weights ...
[strata] loading the experts into RAM (about 40 GB) and locking part of them for the GPU.
         YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES NOW - this is normal.
         Please wait and don't close this window; the browser opens when it is ready.
[strata] still starting (25 s) - please wait ...
[strata] experts loaded: 39.97 GiB at 1.54 GiB/s (40 s so far)
[strata] filling the GPU's expert cache (10977 experts, 17.74 GiB of VRAM) ...
[strata] almost ready ...
Traceback (most recent call last):
  File "/llm/Strata/serve/server.py", line 5015, in <module>
    sys.exit(main())
             ~~~~^^
  File "/llm/Strata/serve/server.py", line 4854, in main
    engine = StrataEngine(exe, engine_args(cfg) + (effort_end or []), cwd=cfg.get("cwd"), log=cfg.get("log"),
                          env=env, lazy=lazy)
  File "/llm/Strata/serve/server.py", line 631, in __init__
    raise RuntimeError("the engine exited before it was ready" + (f" (see {log})" if log else "") +
                       start_failure_hint(log, log_start) + start_log_tail(log, log_start))
RuntimeError: the engine exited before it was ready (see /llm/Strata/strata-swift-iq3_xxs.log)
the engine log's last lines:
  strata generate: loaded 39.97 GiB at 1.54 GiB/s
  strata generate: 1 pool workers on logical processors 2, the host thread on 0 (draining too) (--host-core first)
  strata generate: expert cache auto: 18.42 GiB free, 700 MiB reserved (+0 MiB for the draft head) -> 8176 slots
  strata generate: expert cache 10977 slots, 17.74 GiB of VRAM; policy is
  strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
                   so a reply can differ slightly from a run without the cache (same quality:
                   bench/results/2026-09-27-cache-parity).
                   PROFILE, ranked by routing frequency, no eviction.
  strata generate: pre-filled 10977 of 10977 slots from the profile in 0.7 s (28150 MB/s); slot 0 verified
  strata generate: layer split, CUDA1: 10.32 GiB free of 15.51, room for experts 9.54 GiB
  strata generate: layer split: CUDA1 runs layers 47-47, expert cache 512 slots (1.04 GiB), 512 of its 512 profiled pairs; slot 0 verified
  strata generate: layer split: CUDA0 runs layers 0-46
  strata generate: R4 hit path ON - resident experts are computed on the GPU
  strata generate: 1 expert-pool workers + the host thread
  strata generate: session is up (engine 0.1.40)
  strata generate: token graph hit path: 11489 resident experts, decided on the device
  strata serve: the prompt path allocates its own buffers (too few cache slots to borrow)
  strata serve: prefill: device buffers for a chunk of 1024 tokens do not fit (1 of 24026 MiB free at the failure; 238 MiB taken, the failing buffer wanted 8 MiB): trying a 512-token chunk
  strata serve: prefill: device buffers for a chunk of 512 tokens do not fit (1 of 24026 MiB free at the failure; 167 MiB taken, the failing buffer wanted 8 MiB)
  strata serve: the GPU has too little free VRAM for the prompt path: turn images off (setup: --vision no), close other programs using the GPU, use a shorter context, or read prompts in smaller chunks (--prefill 512)
```

Config generated by `setup.sh`:

```
{
 "exe": "/llm/Strata/engine/strata",
 "args": [
  "--pack",
  "/llm/Strata-data/packs/swift-iq3_xxs",
  "--native",
  "/llm/Strata-data/models/swift-IQ3_XXS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf",
  "--ple-gguf",
  "/llm/Strata-data/models/swift-IQ3_XXS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf",
  "--expert-profile",
  "/llm/Strata/data/expert-profile.bin",
  "--expert-cache",
  "auto",
  "--prefill",
  "auto",
  "--spec",
  "4",
  "--spec-min-p",
  "0.5",
  "--mtp",
  "/llm/Strata-data/mtp/rt",
  "--max-context",
  "524288",
  "--rope-scaling",
  "yarn",
  "--rope-scale",
  "2",
  "--kv",
  "int8",
  "--kv-resident",
  "32768",
  "--remote-expert-opt"
 ],
 "cwd": "/llm/Strata",
 "tokenizer": "/llm/Strata-data/packs/swift-iq3_xxs/tokenizer",
 "model_name": "swift-1.5-iq3_xxs",
 "log": "/llm/Strata/strata-swift-iq3_xxs.log",
 "lib_dirs": [
  "/opt/cuda/bin",
  "/opt/cuda/lib64"
 ],
 "port": 8080,
 "gpu": [
  0,
  1
 ],
 "gpus_asked": true,
 "layer_split": "auto"
}
```

Sur le site

Liens install, modèles, releases.