Issues / #1138
#1138 Windows: higher process/thread priority improves decode throughput under CPU contention, with a measured background-work trade-off
open · @Unmaple · 0 コメント · GitHub で見る
BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows
本文
## Summary On an Intel hybrid-core Windows system, raising Strata's process and thread priorities improved decode throughput and reduced the slowdown caused by concurrent CPU work. Normal-priority background computation retained approximately **94–97% of its standalone throughput**, although its response-time tail increased. These preliminary results suggest that an **opt-in high-priority mode** might be useful. ## Configuration - Intel Core i5-14600K: **6 P-cores + 8 E-cores**, RTX 2080 Ti, Windows. - **13 CPU expert-pool workers plus the calling thread**, without a fixed CPU affinity. - Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S; PCIe fraction 0.14; up to three MTP draft tokens. - Fixed short math prompt: 99 input tokens and 256 generated tokens; greedy generation. - Frozen initial expert profile, with dynamic expert-cache replacement enabled. - All engine settings were identical between groups; only process and thread priorities differed. - Same instrumented executable in both groups, based on f7407d9: common buffered CPU-call timing and resident-RAM controls, without arithmetic-kernel changes. The full cold-expert RAM complement was retained in both groups. Priority conditions: - **N group:** Normal process class; enumerated threads at relative priority 0. - **H group:** High process class; enumerated threads uniformly raised to THREAD_PRIORITY_HIGHEST (+2)., with priorities read back to verify the changes. - Background probe: Normal process class and Normal thread priority throughout. CUDA stream priorities were unchanged. This test did not use Realtime priority. ## Method We tested one and four continuously computing background threads, with **four paired N/H runs per background-load condition**. Run order alternated between N→H and H→N, using fresh engine processes and the same GPU hardware warm-up. Each busy thread performed fixed integer computations. Progress was recorded after each fixed block of 2,097,152 LCG iterations, allowing both wall-clock completion time and throughput to be measured. Thread CPU time was also recorded. A separate Normal-priority periodic probe targeted 100µs intervals. Standalone background baselines were measured before and after the engine comparisons, with the same periodic probe present. All paired generated **token IDs were identical**. Generated-token, MTP accepted/offered, and routing/cache summary counters also matched. ## Decode throughput Values below are total generated tokens divided by total decode time within each condition. | Continuous background threads* | N group | H group | Change | |---|---:|---:|---:| | 1 | 39.98 tok/s | 42.87 tok/s | +7.2% | | 4 | 34.31 tok/s | 41.42 tok/s | +20.7% | *Both conditions also included the periodic probe. Increasing the background load from one to four busy threads reduced decode throughput by approximately **14.2% in the N group**, versus **3.4% in the H group**. ## Background-work throughput Background throughput is normalized to the corresponding standalone baseline, using the average of the before/after baselines. Engine-run throughput is aggregated over the measured decode intervals. | Continuous background threads | N group | H group | |---|---:|---:| | 1 | 98.0% of baseline | 97.3% of baseline | | 4 | 98.2% of baseline | 94.3% of baseline | Relative to N, H reduced busy-thread throughput by approximately **0.7% with one thread** and **4.1% with four threads**. Completing the same amount of work therefore took approximately **0.7% / 4.3% longer**, respectively, using the geometric mean of the paired ratios. The throughput cost was modest in this workload, but responsiveness and throughput should be considered separately. ## Responsiveness and engine call tails | Continuous background threads | Probe P99.9 lateness: N | Probe P99.9 lateness: H | |---|---:|---:| | 1 | 28µs | 193µs | | 4 | 364µs | 1,959µs | These are medians of the per-run P99.9 values, covering executed callbacks. Missed periodic slots are counted separately: with four busy threads, the median skipped-slot fraction increased from **1.28% to 17.86%**. Individual H runs also contained maximum probe delays of approximately **35–93ms**. Conversely, the engine's CPU-pool call P99.9 decreased from approximately **13.56ms to 2.10ms** in the four-thread condition (medians of per-run call percentiles). ## Limitations - One machine, one short-context workload, four independent process pairs per load condition, and 256 output tokens per run. This is preliminary evidence, not a universal performance or responsiveness guarantee. - The 100µs probe remained continuously runnable using QPC and SwitchToThread, adding CPU contention; the same timer-arm overhead remained in all conditions. An initial sleeping-timer calibration did not achieve the required interval. This is **not a direct browser/UI responsiveness test**. - Probe statistics cover an estimated decode interval reconstructed from the request-end timestamp and engine decode duration. CPU-call tail statistics cover all recorded multi-dispatch calls and are not timestamp-aligned to individual probe events. - Priority changes affected both process class and thread-relative priorities; their individual contributions were not isolated. - The pool's default 20ms spin-before-sleep policy remained unchanged. Its contribution to the observed trade-off has not been isolated. Bursty expert computation alone does not establish that idle gaps release the CPU to other programs. - Identical MTP summary counters do not constitute a check of every individual draft sequence. ## Possible upstream direction Would a **user-selectable Windows priority mode** be appropriate, retaining Normal priority as the default and preserving any existing relative thread-priority differences? This is related to the CPU contention discussed in #921, but tests a different variable: **process/thread priority rather than spin duration**.
関連リンク
インストール・モデル・リリースへの站内リンク。