反馈 / #1684
#1684 [Feature]: a repeatable CPU prompt-read share: `auto` changes short-prompt outputs run to run, a fixed share does not
open · @KarlGrier · 0 评论 · 去 GitHub 看
Setup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux
说明
### What do you want to do? Keep v0.1.41's CPU prompt-read share (about 17-21 % faster first tokens on short prompts here) and still get the same output for the same input. That would keep "same request, same answer" checks and conversation resume checks working. With the default (`STRATA_PREFILL_CPU_SHARE` unset, which behaves as `auto` for chunks below 1,024 tokens), the output of a short prompt changes from run to run. `auto` times layers with and without the share and keeps sharing by wall time, so which experts the CPU computes (in its own activation format) depends on timing. A **fixed** share repeated in every case we tried, and here it was as fast as `auto` or faster. Could the default become repeatable, for example `auto` deciding each layer's share once and then keeping it, or a fixed share by default? Or could the docs say that a fixed value repeats and `auto` does not? **Repeatability, stock v0.1.41, 2026-10-09.** `strata generate --stats`, 1,024 tokens after a 136-token math prompt and a 151-token code prompt, greedy and seeded sampling (seed 11, temperature 1.0). The expert cache was held fixed (`--expert-cache 3500 --adapt-swaps 0 --suffix-draft 0`). Two separate engine starts per setting: | `STRATA_PREFILL_CPU_SHARE` | items whose two runs gave the same output | prompt read (ms) | |---|---|---| | unset (`auto`, the default) | **0 of 4** | 406-505 | | `0.4` (fixed) | **4 of 4** | 413-437 | | `0` (off) | 7 of 7 on 2026-10-08 (these four plus 2K / 32K / 112K-token prompts); the same bytes as v0.1.40.2 | 502-545 | On 2026-10-08 two default runs of the same seven items also agreed on only the three long-prompt items, which the share does not touch. **Speed, served, stock v0.1.41, 2026-10-09.** A B C C B A (auto / 0.4 / off), a fresh server per arm. A 9.5K-token and two ~600-token warm-ups, then 6 prompts of 512 and 6 of 1,000 tokens per arm, each a different text, 64 tokens out: | prompt | `auto` TTFT (engine read) | `0.4` | off | |---|---|---|---| | 512 tokens | 579 ms (547) | **576 ms (545)** | 693 ms (664) | | 1,000 tokens | 716 ms (685) | **679 ms (648)** | 864 ms (834) | The two arms of each setting were within 1 %. **What it breaks for us.** With the share on (measured on our local build, which carries a conversation cache), a 96-token follow-up read after 5,726 reused tokens is not repeatable. Two identical direct reads of the same conversation differ in 31 of 1,536 teacher-forced argmax tokens. So a resumed conversation cannot be checked bit for bit against a direct read. With the share off, both are identical. Quality looked unchanged. Teacher-forced NLL on vs off was −0.010 nats on average over six 16K-112K texts read with a 700-token last chunk, about as much as two other rounding-level changes in our build move the same texts, and within ±0.002 nats on two short texts. For now we run with `STRATA_PREFILL_CPU_SHARE=0`. Setup: RTX 5080 16 GB + RTX 4060 Ti 8 GB (helper expert cache, `--expert-cache-device1 3300`), Ryzen 9 9950X3D (16 cores), 96 GB DDR5, Linux 7.0, CUDA 13.3, Qwen3.8-Flash-Next GSQ-RCO IQ3_S, `--prefill auto:16384 --spec 4 --kv int8 --kv-resident 32768`, `STRATA_IO_THREADS=64`. Not tested: whether a fixed share stays repeatable in served sessions with the adaptive tier on, values other than 0.4, other CPUs or GPUs, AMD (HIP), a layer split, and `STRATA_PREFILL_CPU_SHARE_MAX` above 1,024. Related: #1595 (the default's speed on a V100 with every expert in VRAM). Measured and drafted with Claude Code.
本站相关内容
相关页面的快捷入口。