反馈 / #1723

#1723 Strata v0.1.41 - crash unsloth-ud-q4_k_xl

open · @ukrolelo · 0 评论 · 去 GitHub 看

BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows

说明

After few hours of usage unsloth-ud-q4_k_xl this:


[strata] thinking: 701 of max 186510 tokens, 38.4 tok/s, 22 s
[strata] reading the prompt: 30,453 of 30,458 tokens, 3 s so far
[strata] thinking: 701 of max 186510 tokens, 24.8 tok/s, 32 s
[strata] the engine stopped unexpectedly (exit code 3221226505). The engine exited (code 3221226505). Its last log line: strata batch: slot 0 gave back 29772 tokens of this conversation (its turn checkpoint) in 154.9 ms - if that does not explain it, please report it at github.com/Niko1221/Strata/issues with the log. The next request starts the engine again. Its log: A:\0_strata\strata-unsloth-ud-q4_k_xl.log
[strata] the engine stopped unexpectedly (exit code 3221226505). The engine exited (code 3221226505). Its last log line: strata batch: slot 0 gave back 29772 tokens of this conversation (its turn checkpoint) in 154.9 ms - if that does not explain it, please report it at github.com/Niko1221/Strata/issues with the log. The next request starts the engine again. Its log: A:\0_strata\strata-unsloth-ud-q4_k_xl.log
[strata] done: 0 tokens in 18 s (0.0 tok/s) (error, cancel=False)
[strata] done: 701 tokens in 39 s (20.1 tok/s) (error, cancel=False), expert cache 99.8% hit
[strata] the engine had stopped (exit code 3221226505); starting it again (a minute or two) ...
[strata] starting the engine: reading the model's weights ...
[strata] loading the experts into RAM (tens of GB) and locking part of them for the GPU.
         YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES NOW - this is normal.
         Please wait and don't close this window; the browser opens when it is ready.
[strata] still starting (28 s) - please wait ...


strata serve: suffix drafts: 1 windows, 2 of 3 drafts accepted
strata batch: slot 1 takes 154181 tokens (copied in 835.3 ms)
strata serve: prompt 154181 tokens = 154180 reused + 1 read in 2 ms (465.4 tok/s), 1 generated in 12 ms (85.4 tok/s), drafts accepted 0 of 0, 0 checkpoints
strata serve: decode expert cache hit rate: 99.8% (479 hits / 480 lookups)
strata serve: KV streaming: 90.62% of 671676 block reads hit VRAM, 253.7 MiB read from RAM
strata serve: conversation cache: parked 154181 tokens in 578.2 ms; parked=6 bytes=10274268092 evictions=12 snapshot_bytes=2472151580 reused_kv_bytes=0
strata batch: slot 0 gave back 29772 tokens of this conversation (its turn checkpoint) in 154.9 ms
[strata] 2026-10-09 16:05:52 engine started: A:\0_strata\engine\strata.exe --serve --pack C:\Strata-data\packs\unsloth-ud-q4_k_xl --native C:\Strata-data\models\unsloth-UD-Q4_K_XL\Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf --expert-profile A:\0_strata\data\expert-profile.bin --expert-cache auto --prefill auto --spec 4 --mtp A:\Strata-data\mtp\rt --max-context 340000 --kv int8 --kv-resident 32768 --vision --vram-reserve-mib 700 --pcie-frac 0.10 --spec-min-p 0.20 --conversation-cache-mib 32768 --conversation-cache-slots 6 --prompt-cache 6 --batch 2
strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
strata generate: native pack: C:\Strata-data\packs\unsloth-ud-q4_k_xl experts (largest blob 3.99 MB), token embedding Q8_0 in mapped host memory (644 MiB, 1.1 s)
strata generate: 1416 MiB of weights loaded from C:\Strata-data\packs\unsloth-ud-q4_k_xl in 0.3 s (303 canonical tensors skipped: served natively)
strata generate: 301 native projection matrices, 2975.00 MiB of weights, in 4.8 s
strata generate: the SSD is kept awake while rows are read: one page of the table after 100 ms without a read, until 60 s after the last request (STRATA_SSD_KEEPALIVE=0 turns it off)
strata generate: profile A:\0_strata\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata generate: KV streaming: 32768 of 340000 cells per QSA layer in VRAM, the K/V in 4.01 GiB of pinned RAM
strata generate: PLE on, table 320001536 rows of C:\Strata-data\models\unsloth-UD-Q4_K_XL\Qwen3.8-Flash-Next-UD-Q4_K_XL-00002-of-00004.gguf
strata generate: --batch: 2 slot sessions on CUDA0 (1.08 GiB each); 38.67 GiB free
strata generate: --batch 2: the slot sessions take 2.16 GiB of VRAM on CUDA0 that the expert cache would otherwise hold
strata mtp: draft layer loaded, 835 MiB of VRAM (experts 675, dense 111), files read in 0.47 s (1675 MiB/s)
strata generate: experimental native Q5_K head, 675430400 bytes, in 0.9 s
strata generate: GPU 0: NVIDIA RTX PRO 5000 Blackwell, compute capability 12.0
strata generate: expert arena read unbuffered (0 of 48 probe reads from the file cache; 102.3 GiB available, 103.7 GiB of files)
strata generate: expert arena: locked 3 MiB via working-set minimum + VirtualLock; cudaHostRegister of the whole arena FAILED (out of memory); 48 slices pinned (71 GiB); large pages refused for 77022101504 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata generate: loaded 71.73 GiB at 1.83 GiB/s
strata generate: CPU pool tasks/phase: 72 (automatic), participating threads: 24
strata generate: 23 pool workers on logical processors 2,4,6,8,10,12,14,16,18,20,22,24,26,28,30,32,34,36,38,40,42,44,46, the host thread on 0 (draining too) (--host-core first)
strata generate: expert cache auto: 37.15 GiB free, 700 MiB reserved (+281 MiB for the draft head) -> 9729 slots
strata generate: expert cache 12389 slots, 36.18 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
                 so a reply can differ slightly from a run without the cache (same quality:
                 bench/results/2026-09-27-cache-parity).
                 PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 12389 of 12389 slots from the profile in 1.5 s (25595 MB/s); slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 23 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.41)
strata generate: token graph hit path: 12389 resident experts, decided on the device
strata serve: prompt chunk auto: 8192 tokens, a 384-slot ring
strata serve: the prompt path borrows 1663 CUDA0 cache slots (4.87 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token 

本站相关内容

相关页面的快捷入口。