Issues / #511

#511 Engine stopped unexpectedly

closed · @sinand99 · 1 commentaires · Sur GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows

Description

Stopped abruptly when working via OMP harness.

Log
---------
strata-vision: on the CPU, 12 threads, no warm-up
strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
strata generate: native pack: G:\AI\Text\Apps\Strata-data\packs\iq3_s experts (largest blob 2.66 MB), token embedding IQ4_XS in mapped host memory (322 MiB)
strata generate: 1466 MiB of weights loaded from G:\AI\Text\Apps\Strata-data\packs\iq3_s (302 canonical tensors skipped: served natively)
strata generate: 300 native projection matrices, 2018.88 MiB of weights
strata generate: the SSD is kept awake while rows are read: one page of the table after 100 ms without a read, until 60 s after the last request (STRATA_SSD_KEEPALIVE=0 turns it off)
strata generate: profile G:\AI\Text\Apps\Strata\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata generate: KV streaming: 32768 of 98304 cells per QSA layer in VRAM, the K/V in 1.16 GiB of pinned RAM
strata generate: PLE on, table 320001536 rows of G:\AI\Text\Apps\Strata-data\models\IQ3_S\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf
strata mtp: draft layer loaded, 834 MiB of VRAM (experts 675, dense 111), files read in 0.40 s (1945 MiB/s)
strata generate: GPU 0: NVIDIA GeForce RTX 3070 Ti, compute capability 8.6
strata generate: expert arena: locked 2365 MiB via working-set minimum + VirtualLock; cudaHostRegister of the whole arena FAILED (out of memory); 46 slices pinned (44 GiB); large pages refused for 50295996416 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata generate: loaded 46.84 GiB at 3.62 GiB/s
strata generate: experimental native Q5_K head, 521472000 bytes
strata generate: expert cache auto: 1.18 GiB free, 150 MiB reserved (+85 MiB for the draft head) -> 382 slots
strata generate: expert cache 487 slots, 0.95 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
                 so a reply can differ slightly from a run without the cache (same quality:
                 bench/results/2026-09-27-cache-parity).
                 PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 487 of 487 slots from the profile; slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 11 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.35)
strata generate: token graph hit path: 487 resident experts, decided on the device
strata serve: prompt chunk auto: 512 tokens
strata serve: the prompt path borrows 250 CUDA0 cache slots (0.49 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token and stream, the down projection's activations staged ahead by cp.async); checked bit for bit against the plain read on this card (STRATA_HC_SPLIT=1 or 0 for the earlier ones)
strata verify: window up to 4 tokens, 63.2 MiB of device buffers
strata mtp: draft head over 40525 tokens (81.2 MiB)
strata serve: 0 MiB of VRAM free with everything loaded - LOW: requests may stall; add --vram-reserve-mib 662 to the config's args (or lower --max-context)
strata verify: captured the 4-token window (upload no error, sync no error)
strata verify: captured the 1-token window (upload no error, sync no error)
strata verify: captured the 2-token window (upload no error, sync no error)
strata serve: prompt 145 tokens = 0 reused + 145 read in 1794 ms (80.8 tok/s), 65 generated in 2138 ms (30.4 tok/s), drafts accepted 27 of 30, 1 checkpoints
strata serve: decode expert cache hit rate: 24.3% (6228 hits / 25642 lookups)
strata serve: KV streaming: 98.33% of 38184 block reads hit VRAM, 2.6 MiB read from RAM
strata serve: suffix drafts: 1 windows, 1 of 3 drafts accepted
strata serve: prompt 9114 tokens = 0 reused + 9114 read in 37842 ms (240.8 tok/s), 89 generated in 3711 ms (24.0 tok/s), drafts accepted 44 of 51, 2 checkpoints
strata serve: decode expert cache hit rate: 26.0% (9585 hits / 36896 lookups)
strata serve: KV streaming: 96.25% of 622356 block reads hit VRAM, 94.1 MiB read from RAM
strata serve: suffix drafts: 5 windows, 10 of 15 drafts accepted
strata verify: captured the 3-token window (upload no error, sync no error)
strata serve: prompt 9223 tokens = 9203 reused + 20 read in 626 ms (32.0 tok/s), 75 generated in 3482 ms (21.5 tok/s), drafts accepted 39 of 45, 3 checkpoints
strata serve: decode expert cache hit rate: 24.5% (7598 hits / 30973 lookups)
strata serve: KV streaming: 97.98% of 1244712 block reads hit VRAM, 101.4 MiB read from RAM
strata serve: suffix drafts: 8 windows, 17 of 19 drafts accepted
strata serve: prompt 10783 tokens = 9298 reused + 1485 read in 6215 ms (238.9 tok/s), 24861 generated in 1027305 ms (24.2 tok/s), drafts accepted 9651 of 10829, 4 checkpoints
strata serve: decode expert cache hit rate: 33.3% (3353029 hits / 10076167 lookups)
strata serve: KV streaming: 99.93% of 161727804 block reads hit VRAM, 478.6 MiB read from RAM
strata serve: suffix drafts: 189 windows, 168 of 242 drafts accepted
strata serve: prompt 680 tokens = 0 reused + 680 read in 3401 ms (200.0 tok/s), 82 generated in 2835 ms (28.9 tok/s), drafts accepted 34 of 38, 1 checkpoints
strata serve: decode expert cache hit rate: 24.2% (7874 hits / 32475 lookups)
strata serve: KV streaming: 98.82% of 194412 block reads hit VRAM, 9.2 MiB read from RAM
strata serve: prompt 35827 tokens = 0 reused + 35827 read in 137254 ms (261.0 tok/s), 16172 generated in 554372 ms (29.2 tok/s), drafts accepted 7946 of 8232, 4 checkpoints
strata serve: decode expert cache hit rate: 29.4% (1840292 hits / 6265663 lookups)
strata serve: KV streaming: 99.50% of 101438772 block reads hit VRAM, 2024.0 MiB read from RAM
strata serve: suffix drafts: 235 windows, 607 of 661 drafts accepted
strata serve: prompt 52035 tokens = 51998 reused + 37 read in 1145 ms (32.3 tok/s), 292 generated in 12190 ms (24.0 tok/s), drafts accepted 128 of 139, 5 checkpoints
strata serve: decode expert cache hit rate: 23.9% (27254 hits / 114160 lookups)
strata serve: KV streaming: 99.48% of 103527684 block reads hit VRAM, 2179.7 MiB read from RAM
strata serve: suffix drafts: 9 windows, 27 of 27 drafts accepted
strata serve: prompt 52507 tokens = 52326 reused + 181 read in 1759 ms (102.9 tok/s), 393 generated in 15476 ms (25.4 tok/s), drafts accepted 197 of 209, 6 checkpoints
strata serve: decode expert cache hit rate: 28.3% (43681 hits / 154532 lookups)
strata serve: KV streaming: 99.46% of 106054116 block reads hit VRAM, 2295.7 MiB read from RAM
strata serve: suffix drafts: 28 windows, 81 of 84 drafts accepted
strata serve: prompt 54908 tokens = 52900 reused + 2008 read in 9565 ms (209.9 tok/s), 133 generated in 5409 ms (24.6 tok/s), drafts accepted 59 of 62, 6 checkpoints
strata serve: decode expert cache hit rate: 25.1% (13024 hits / 51856 lookups)
strata serve: KV streaming: 99.46% of 106922964 block reads hit VRAM, 2326.5 MiB read from RAM
strata serve: suffix drafts: 3 windows, 7 of 7 drafts accepted
strata serve: prompt 55691 tokens = 55041 reused + 650 read in 4053 ms (160.4 tok/s), 184 generated in 7275 ms (25.3 tok/s), drafts accepted 90 of 93, 6 checkpoints
strata serve: decode expert cache hit rate: 24.2% (17017 hits / 70383 lookups)
strata serve: KV streaming: 99.46% of 108099888 block reads hit VRAM, 2370.6 MiB read from RAM
strata serve: suffix drafts: 6 windows, 12 of 14 drafts accepted
strata serve: prompt 56079 tokens = 55874 reused + 205 read in 1781 ms (115.1 tok/s), 114 generated in 4716 ms (24.2 tok/s), drafts accepted 57 of 61, 6 checkpoints
strata serve: decode expert cache hit rate: 22.4% (9921 hits / 44382 lookups)
strata serve: KV streaming: 99.45% of 108851616 block reads hit VRAM, 2401.7 MiB read from RAM
strata serve: suffix drafts: 6 windows, 14 of 16 drafts accepted
strata serve: prompt 56503 tokens = 56192 reused + 311 read in 2054 ms (151.4 tok/s), 108 generated in 4383 ms (24.6 tok/s), drafts accepted 51 of 56, 6 checkpoints
strata serve: decode expert cache hit rate: 22.3% (9477 hits / 42446 lookups)
strata serve: KV streaming: 99.45% of 109572564 block reads hit VRAM, 2427.0 MiB read from RAM
strata serve: suffix drafts: 5 windows, 9 of 13 drafts accepted
strata serve: prompt 56679 tokens = 56610 reused + 69 read in 2178 ms (31.7 tok/s), 119 generated in 4786 ms (24.9 tok/s), drafts accepted 59 of 59, 6 checkpoints
strata serve: decode expert cache hit rate: 25.2% (11404 hits / 45244 lookups)
strata serve: KV streaming: 99.45% of 110731020 block reads hit VRAM, 2451.2 MiB read from RAM
strata serve: suffix drafts: 5 windows, 15 of 15 drafts accepted
strata serve: prompt 56976 tokens = 56798 reused + 178 read in 1656 ms (107.5 tok/s), 82 generated in 3397 ms (24.1 tok/s), drafts accepted 39 of 39, 6 checkpoints
strata serve: decode expert cache hit rate: 23.5% (7338 hits / 31280 lookups)
strata serve: KV streaming: 99.45% of 111267120 block reads hit VRAM, 2477.2 MiB read from RAM
strata serve: prompt 59648 tokens = 57058 reused + 2590 read in 11943 ms (216.9 tok/s), 254 generated in 10249 ms (24.8 tok/s), drafts accepted 112 of 121, 6 checkpoints
strata serve: decode expert cache hit rate: 31.2% (31721 hits / 101631 lookups)
strata serve: KV streaming: 99.43% of 112912392 block reads hit VRAM, 2573.0 MiB read from RAM
strata serve: suffix drafts: 1 windows, 1 of 1 drafts accepted
strata serve: prompt 60204 tokens = 59901 reused + 303 read in 1884 ms (160.9 tok/s), 193 generated in 7702 ms (25.1 tok/s), drafts accepted 86 of 93, 6 checkpoints
strata serve: decode expert cache hit rate: 30.9% (23929 hits / 77319 lookups)
strata serve: KV streaming: 99.43% of 114175596 block reads hit VRAM, 2601.3 MiB read from RAM
strata serve: prompt 60465 tokens = 60397 reused + 68 read in 2150 ms (31.6 tok/s), 234 generated in 9170 ms (25.5 tok/s), drafts accepted 115 of 122, 6 checkpoints
strata serve: decode expert cache hit rate: 28.5% (26123 hits / 91815 lookups)
strata serve: KV streaming: 99.44% of 116073492 block reads hit VRAM, 2633.8 MiB read from RAM
strata serve: suffix drafts: 9 windows, 23 of 27 drafts accepted
strata serve: verify: timed out at layer 31; its GPU waits were released but the GPU did not finish within 5 s (#267)


On the second run, another stall happened:

Log
-----

strata serve: no progress for 60 s during a request (decode -1) - stopping the engine so the server starts it again (issue #29)
strata serve: stall report (engine 0.1.35): stage "decode -1" for 61 s; 0 layers served since the last finished step (0 = stopped, more = slow)
  expert pool: epoch 443080, batch epoch 443080: 36 of 36 jobs claimed, 36 done; 11 of 11 workers parked, 11 sleeping; mode 0
  expert pool threads: w0=sleeping w1=sleeping w2=sleeping w3=sleeping w4=sleeping w5=sleeping w6=sleeping w7=sleeping w8=sleeping w9=sleeping w10=sleeping; host idle for 61499 ms
  verify window (last window, not the current stage): 2 tokens at position 73860, host at layer step 48; the GPU rang 48; flags: served 48, plan (A) 48, copies (B) 48
  memory: 50191 MiB resident, 59875 MiB committed, 1534 MiB RAM available; 25758395 page faults so far
  2 s later:
  expert pool: epoch 443080, batch epoch 443080: 36 of 36 jobs claimed, 36 done; 11 of 11 workers parked, 11 sleeping; mode 0
  expert pool threads: w0=sleeping w1=sleeping w2=sleeping w3=sleeping w4=sleeping w5=sleeping w6=sleeping w7=sleeping w8=sleeping w9=sleeping w10=sleeping; host idle for 63508 ms
  verify window (last window, not the current stage): 2 tokens at position 73860, host at layer step 48; the GPU rang 48; flags: served 48, plan (A) 48, copies (B) 48
  memory: 50191 MiB resident, 59875 MiB committed, 1537 MiB RAM available; 25758396 page faults so far
  wrote the thread stacks to G:\AI\Text\Apps\Strata\strata-stall-9804.dmp (attach it to the issue)
strata: released the verify window's GPU waits (#267): the GPU did not finish within 5 s

[strata-stall-9804.dmp](https://gith

Sur le site

Liens install, modèles, releases.