Issues / #779
#779 laptop crash with latest
open · @XeonG · 3 コメント · GitHub で見る
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
本文
<img width="917" height="501" alt="Image" src="https://github.com/user-attachments/assets/c7bf527c-19c1-45f3-ba7e-80e6b67a4d8d" />
win11, 16vram, 64gb ram usually 10gb free ram still with this running, vram is a different story.
```
strata generate: PCIe probe: 24.0 GB/s host->device (best of 24.0 24.0 24.0 24.0) -> pcie_frac 0.55 (default 0.55)
strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
strata generate: native pack: E:\Strata-data\packs\swift-iq3_xxs experts (largest blob 2.33 MB), token embedding IQ3_S in mapped host memory (260 MiB)
strata generate: 1466 MiB of weights loaded from E:\Strata-data\packs\swift-iq3_xxs (302 canonical tensors skipped: served natively)
strata generate: 300 native projection matrices, 1781.11 MiB of weights
strata generate: the SSD is kept awake while rows are read: one page of the table after 100 ms without a read, until 60 s after the last request (STRATA_SSD_KEEPALIVE=0 turns it off)
strata generate: profile E:\Strata\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata generate: KV streaming: 32768 of 131072 cells per QSA layer in VRAM, the K/V in 1.55 GiB of pinned RAM
strata generate: PLE on, table 320001536 rows of E:\Strata-data\models\swift-IQ3_XXS\Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf
strata mtp: draft layer loaded, 835 MiB of VRAM (experts 675, dense 111), files read in 0.20 s (3885 MiB/s)
strata generate: experimental native Q5_K head, 437043200 bytes
strata generate: GPU 0: NVIDIA GeForce RTX 4090 Laptop GPU, compute capability 8.9
strata generate: expert arena read unbuffered (2 of 32 probe reads from the file cache; 14.5 GiB available, 70.7 GiB of files)
strata generate: expert arena: VirtualLock stopped at 0 of 11452 MiB (error 87); cudaHostRegister of the whole arena FAILED (out of memory); 36 slices pinned (28 GiB); large pages (2097152 B)
strata generate: loaded 39.97 GiB at 1.42 GiB/s
strata generate: expert cache auto: 9.36 GiB free, 700 MiB reserved (+184 MiB for the draft head) -> 3915 slots
strata generate: expert cache 5235 slots, 8.49 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
so a reply can differ slightly from a run without the cache (same quality:
bench/results/2026-09-27-cache-parity).
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 5235 of 5235 slots from the profile; slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 23 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.39)
strata generate: token graph hit path: 5235 resident experts, decided on the device
strata serve: prompt chunk auto: 8192 tokens, a 96-slot ring
strata serve: the prompt path borrows 2178 CUDA0 cache slots (3.50 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token and stream, the down projection's activations staged ahead by cp.async); checked bit for bit against the plain read on this card (STRATA_HC_SPLIT=1 or 0 for the earlier ones)
strata verify: window up to 6 tokens, 69.7 MiB of device buffers
strata mtp: draft head over 106299 tokens (178.4 MiB)
strata serve: 444 MiB of VRAM free with everything loaded
strata generate: PCIe probe: 23.9 GB/s host->device (best of 23.9 23.7 23.9 23.9) -> pcie_frac 0.55 (default 0.55)
strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
strata generate: native pack: E:\Strata-data\packs\swift-iq3_xxs experts (largest blob 2.33 MB), token embedding IQ3_S in mapped host memory (260 MiB)
strata generate: 1466 MiB of weights loaded from E:\Strata-data\packs\swift-iq3_xxs (302 canonical tensors skipped: served natively)
strata generate: 300 native projection matrices, 1781.11 MiB of weights
strata generate: the SSD is kept awake while rows are read: one page of the table after 100 ms without a read, until 60 s after the last request (STRATA_SSD_KEEPALIVE=0 turns it off)
strata generate: profile E:\Strata\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata generate: KV streaming: 32768 of 131072 cells per QSA layer in VRAM, the K/V in 1.55 GiB of pinned RAM
strata generate: PLE on, table 320001536 rows of E:\Strata-data\models\swift-IQ3_XXS\Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf
strata mtp: draft layer loaded, 835 MiB of VRAM (experts 675, dense 111), files read in 0.69 s (1134 MiB/s)
strata generate: experimental native Q5_K head, 437043200 bytes
strata generate: GPU 0: NVIDIA GeForce RTX 4090 Laptop GPU, compute capability 8.9
strata generate: expert arena read unbuffered (1 of 32 probe reads from the file cache; 14.9 GiB available, 70.7 GiB of files)
strata generate: expert arena: VirtualLock stopped at 0 of 11452 MiB (error 87); cudaHostRegister of the whole arena FAILED (out of memory); 36 slices pinned (28 GiB); large pages (2097152 B)
strata generate: loaded 39.97 GiB at 1.43 GiB/s
strata generate: expert cache auto: 9.36 GiB free, 700 MiB reserved (+184 MiB for the draft head) -> 3915 slots
strata generate: expert cache 5235 slots, 8.49 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
so a reply can differ slightly from a run without the cache (same quality:
bench/results/2026-09-27-cache-parity).
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 5235 of 5235 slots from the profile; slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 23 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.39)
strata generate: token graph hit path: 5235 resident experts, decided on the device
strata serve: prompt chunk auto: 8192 tokens, a 96-slot ring
strata serve: the prompt path borrows 2178 CUDA0 cache slots (3.50 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token and stream, the down projection's activations staged ahead by cp.async); checked bit for bit against the plain read on this card (STRATA_HC_SPLIT=1 or 0 for the earlier ones)
strata verify: window up to 6 tokens, 69.7 MiB of device buffers
strata mtp: draft head over 106299 tokens (178.4 MiB)
strata serve: 444 MiB of VRAM free with everything loaded
strata generate: PCIe probe: 24.0 GB/s host->device (best of 23.8 23.9 23.9 24.0) -> pcie_frac 0.55 (default 0.55)
strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
strata generate: native pack: E:\Strata-data\packs\swift-iq3_xxs experts (largest blob 2.33 MB), token embedding IQ3_S in mapped host memory (260 MiB)
strata generate: 1466 MiB of weights loaded from E:\Strata-data\packs\swift-iq3_xxs (302 canonical tensors skipped: served natively)
strata generate: 300 native projection matrices, 1781.11 MiB of weights
strata generate: the SSD is kept awake while rows are read: one page of the table after 100 ms without a read, until 60 s after the last request (STRATA_SSD_KEEPALIVE=0 turns it off)
strata generate: profile E:\Strata\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata generate: KV streaming: 32768 of 131072 cells per QSA layer in VRAM, the K/V in 1.55 GiB of pinned RAM
strata generate: PLE on, table 320001536 rows of E:\Strata-data\models\swift-IQ3_XXS\Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf
strata mtp: draft layer loaded, 835 MiB of VRAM (experts 675, dense 111), files read in 0.20 s (3878 MiB/s)
strata generate: experimental native Q5_K head, 437043200 bytes
strata generate: GPU 0: NVIDIA GeForce RTX 4090 Laptop GPU, compute capability 8.9
strata generate: expert arena read unbuffered (0 of 32 probe reads from the file cache; 14.8 GiB available, 70.7 GiB of files)
strata generate: expert arena: VirtualLock stopped at 0 of 11452 MiB (error 87); cudaHostRegister of the whole arena FAILED (out of memory); 36 slices pinned (28 GiB); large pages (2097152 B)
strata generate: loaded 39.97 GiB at 1.43 GiB/s
strata generate: expert cache auto: 9.36 GiB free, 700 MiB reserved (+184 MiB for the draft head) -> 3915 slots
strata generate: expert cache 5235 slots, 8.49 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
so a reply can differ slightly from a run without the cache (same quality:
bench/results/2026-09-27-cache-parity).
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 5235 of 5235 slots from the profile; slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 23 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.39)
strata generate: token graph hit path: 5235 resident experts, decided on the device
strata serve: prompt chunk auto: 8192 tokens, a 96-slot ring
strata serve: the prompt path borrows 2178 CUDA0 cache slots (3.50 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token and stream, the down projection's activations staged ahead by cp.async); checked bit for bit against the plain read on this card (STRATA_HC_SPLIT=1 or 0 for the earlier ones)
strata verify: window up to 6 tokens, 69.7 MiB of device buffers
strata mtp: draft head over 106299 tokens (178.4 MiB)
strata serve: 444 MiB of VRAM free with everything loaded
strata verify: captured the 6-token window (upload no error, sync no error)
strata verify: captured the 4-token window (upload no error, sync no error)
strata verify: captured the 1-token window (upload no error, sync no error)
strata verify: captured the 2-token window (upload no error, sync no error)
strata serve: prompt 53 tokens = 0 reused + 53 read in 895 ms (59.2 tok/s), 33 generated in 603 ms (54.8 tok/s), drafts accepted 20 of 26, 1 checkpoints
strata serve: decode expert cache hit rate: 66.3% (10620 hits / 16022 lookups); 2698 more read by the GPU over PCIe or from another GPU (14.4% of all 18720 routed)
strata serve: KV streaming: 97.90% of 12576 block reads hit VRAM, 1.1 MiB read from RAM
strata verify: captured the 3-token window (upload no error, sync no error)
strata serve: prompt 59 tokens = 48 reused + 11 read in 477 ms (23.0 tok/s), 57 generated in 1053 ms (54.1 tok/s), drafts accepted 31 of 47, 2 checkpoints
strata serve: decode expert cache hit rate: 79.2% (25339 hits / 31976 lookups); 3544 more read by the GPU over PCIe or from another GPU (10.0% of all 35520 routed)
strata serve: KV streaming: 98.98% of 34212 block reads hit VRAM, 1.4 MiB read from RAM
strata serve: prompt 306 tokens = 0 reused + 306 read in 1645 ms (186.0 tok/s), 334 generated in 5313 ms (62.9 tok/s), drafts accepted 213 of 271, 1 checkpoints
strata serve: decode expert cache hit rate: 75.6% (124964 hits / 165201 lookups); 23439 more read by the GPU over PCIe or from another GPU (12.4% of all 188640 routed)
strata serve: KV streaming: 99.65% of 552024 block reads hit VRAM, 7.7 MiB read from RAM
strata serve: suffix drafts: 3 windows, 6 of 9 drafts accepted
strata serve: prompt 306 tokens = 0 reused + 306 read in 1020 ms (300.1 tok/s), 289 generated in 4588 ms (63.0 tok/s), drafts accepted 177 of 223, 1 checkpoints
strata serve: decode expert cache hit rate: 80.6% (116757 hits / 144809 lookups); 16471 more read by the GPU over PCIe or from another GPU (10.2% of all 161280 routed)
strata serve: KV streaming: 99.61% of 453240 block reads hit VRAM, 7.2 MiB read from RAM
strata serve: suffix drafts: 2 windows, 1 of 3 drafts accepted
strata verify: captured the 5-token window (upload no error, sync no error)
strata serve: prompt 28034 tokens = 0 reused + 28034 read in 17605 ms (関連リンク
インストール・モデル・リリースへの站内リンク。