Issues / #535
#535 Thinking speed suddenly drops to 0.5 token/s
closed · @sinand99 · 2 Kommentare · Auf GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
Beschreibung
<img width="884" height="768" alt="Image" src="https://github.com/user-attachments/assets/f6ad390f-1834-4cf6-8066-26822d7da655" />
After some tool calls, the thinking speed suddenly drops to 0.5 token/s and the engines uses only 1 CPU core. Only way to fix it is to restart the engine.
Log:
----------
strata-vision: on the CPU, 12 threads, no warm-up
strata generate: this CPU has no AVX-512: the expert kernels run on AVX-2 (multi-token for the i-quant gate/up rows)
strata generate: native pack: G:\AI\Text\Apps\Strata-data\packs\iq3_s experts (largest blob 2.66 MB), token embedding IQ4_XS in mapped host memory (322 MiB)
strata generate: 1466 MiB of weights loaded from G:\AI\Text\Apps\Strata-data\packs\iq3_s (302 canonical tensors skipped: served natively)
strata generate: 300 native projection matrices, 2018.88 MiB of weights
strata generate: the SSD is kept awake while rows are read: one page of the table after 100 ms without a read, until 60 s after the last request (STRATA_SSD_KEEPALIVE=0 turns it off)
strata generate: profile G:\AI\Text\Apps\Strata\data\expert-profile.bin: 24576 ranked pairs, built for 24576 slots
strata generate: KV streaming: 32768 of 98304 cells per QSA layer in VRAM, the K/V in 1.16 GiB of pinned RAM
strata generate: PLE on, table 320001536 rows of G:\AI\Text\Apps\Strata-data\models\IQ3_S\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf
strata mtp: draft layer loaded, 834 MiB of VRAM (experts 675, dense 111), files read in 0.40 s (1945 MiB/s)
strata generate: GPU 0: NVIDIA GeForce RTX 3070 Ti, compute capability 8.6
strata generate: expert arena: locked 2365 MiB via working-set minimum + VirtualLock; cudaHostRegister of the whole arena FAILED (out of memory); 46 slices pinned (44 GiB); large pages refused for 50295996416 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata generate: loaded 46.84 GiB at 3.30 GiB/s
strata generate: experimental native Q5_K head, 521472000 bytes
strata generate: expert cache auto: 1.00 GiB free, 128 MiB reserved (+85 MiB for the draft head) -> 319 slots
strata generate: expert cache 408 slots, 0.79 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
so a reply can differ slightly from a run without the cache (same quality:
bench/results/2026-09-27-cache-parity).
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 408 of 408 slots from the profile; slot 0 verified
strata generate: R4 hit path ON - resident experts are computed on the GPU
strata generate: 11 expert-pool workers + the host thread
strata generate: session is up (engine 0.1.36)
strata generate: token graph hit path: 408 resident experts, decided on the device
strata serve: prompt chunk 2048 -> 512 tokens so its buffers fit in every expert cache
strata serve: the prompt path borrows 250 CUDA0 cache slots (0.48 GiB)
strata hc: CUDA0: the hyper-connection read runs as staged (the norm per token and stream, the down projection's activations staged ahead by cp.async); checked bit for bit against the plain read on this card (STRATA_HC_SPLIT=1 or 0 for the earlier ones)
strata verify: window up to 4 tokens, 63.2 MiB of device buffers
strata mtp: draft head over 40525 tokens (81.2 MiB)
strata serve: 0 MiB of VRAM free with everything loaded - LOW: requests may stall; add --vram-reserve-mib 640 to the config's args (or lower --max-context)
strata verify: captured the 4-token window (upload no error, sync no error)
strata verify: captured the 1-token window (upload no error, sync no error)
strata verify: captured the 2-token window (upload no error, sync no error)
strata verify: captured the 3-token window (upload no error, sync no error)
strata serve: prompt 38373 tokens = 0 reused + 38373 read in 160723 ms (238.8 tok/s), 6426 generated in 462020 ms (13.9 tok/s), drafts accepted 2753 of 3026, 3 checkpoints
strata serve: decode expert cache hit rate: 24.9% (629117 hits / 2523507 lookups)
strata serve: KV streaming: 99.33% of 41309976 block reads hit VRAM, 1109.1 MiB read from RAM
strata serve: suffix drafts: 99 windows, 173 of 225 drafts accepted
strata serve: prompt 6565 tokens = 0 reused + 6565 read in 21735 ms (302.0 tok/s), 978 generated in 32828 ms (29.8 tok/s), drafts accepted 394 of 420, 1 checkpoints
strata serve: decode expert cache hit rate: 22.2% (83340 hits / 374639 lookups)
strata serve: KV streaming: 99.64% of 6211260 block reads hit VRAM, 91.2 MiB read from RAM
strata serve: suffix drafts: 7 windows, 5 of 7 drafts accepted
strata serve: prompt 38373 tokens = 0 reused + 38373 read in 135239 ms (283.7 tok/s), 5486 generated in 163234 ms (33.6 tok/s), drafts accepted 2321 of 2512, 3 checkpoints
strata serve: decode expert cache hit rate: 24.0% (511432 hits / 2134558 lookups)
strata serve: KV streaming: 99.12% of 35012580 block reads hit VRAM, 1247.6 MiB read from RAM
strata serve: suffix drafts: 49 windows, 36 of 49 drafts accepted
strata serve: prompt 12301 tokens = 0 reused + 12301 read in 43645 ms (281.8 tok/s), 306 generated in 9454 ms (32.4 tok/s), drafts accepted 111 of 130, 1 checkpoints
strata serve: decode expert cache hit rate: 22.4% (27211 hits / 121516 lookups)
strata serve: KV streaming: 98.40% of 2027340 block reads hit VRAM, 131.0 MiB read from RAM
strata serve: suffix drafts: 1 windows, 0 of 1 drafts accepted
strata serve: prompt 21645 tokens = 0 reused + 21645 read in 78475 ms (275.8 tok/s), 1427 generated in 46587 ms (30.6 tok/s), drafts accepted 531 of 580, 3 checkpoints
strata serve: decode expert cache hit rate: 28.4% (159582 hits / 562137 lookups)
strata serve: KV streaming: 99.35% of 9125916 block reads hit VRAM, 240.0 MiB read from RAM
strata serve: suffix drafts: 9 windows, 7 of 9 drafts accepted
strata serve: prompt 23336 tokens = 23072 reused + 264 read in 1610 ms (164.0 tok/s), 205 generated in 6441 ms (31.8 tok/s), drafts accepted 99 of 102, 4 checkpoints
strata serve: decode expert cache hit rate: 22.4% (17455 hits / 78079 lookups)
strata serve: KV streaming: 99.39% of 10438428 block reads hit VRAM, 254.8 MiB read from RAM
strata serve: suffix drafts: 2 windows, 2 of 2 drafts accepted
strata serve: prompt 23993 tokens = 23541 reused + 452 read in 66721 ms (6.8 tok/s), 555 generated in 960774 ms (0.6 tok/s), drafts accepted 197 of 250, 5 checkpoints
strata serve: decode expert cache hit rate: 30.3% (70521 hits / 232674 lookups)
strata serve: KV streaming: 99.52% of 14209572 block reads hit VRAM, 273.8 MiB read from RAMMehr auf der Site
Links zu Install, Modellen, Releases.