Issues / #1138

#1138 Windows: higher process/thread priority improves decode throughput under CPU contention, with a measured background-work trade-off

open · @Unmaple · 0 comments · View on GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsWindows

Description

## Summary

On an Intel hybrid-core Windows system, raising Strata's process and thread priorities improved decode throughput and reduced the slowdown caused by concurrent CPU work.

Normal-priority background computation retained approximately **94–97% of its standalone throughput**, although its response-time tail increased.

These preliminary results suggest that an **opt-in high-priority mode** might be useful. 

## Configuration

- Intel Core i5-14600K: **6 P-cores + 8 E-cores**, RTX 2080 Ti, Windows.
- **13 CPU expert-pool workers plus the calling thread**, without a fixed CPU affinity.
- Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S; PCIe fraction 0.14; up to three MTP draft tokens.
- Fixed short math prompt: 99 input tokens and 256 generated tokens; greedy generation.
- Frozen initial expert profile, with dynamic expert-cache replacement enabled.
- All engine settings were identical between groups; only process and thread priorities differed.
- Same instrumented executable in both groups, based on f7407d9: common buffered CPU-call timing and resident-RAM controls, without arithmetic-kernel changes. The full cold-expert RAM complement was retained in both groups.

Priority conditions:

- **N group:** Normal process class; enumerated threads at relative priority 0.
- **H group:** High process class; enumerated threads uniformly raised to THREAD_PRIORITY_HIGHEST (+2)., with priorities read back to verify the changes.
- Background probe: Normal process class and Normal thread priority throughout.

CUDA stream priorities were unchanged. This test did not use Realtime priority.

## Method

We tested one and four continuously computing background threads, with **four paired N/H runs per background-load condition**. Run order alternated between N→H and H→N, using fresh engine processes and the same GPU hardware warm-up.

Each busy thread performed fixed integer computations. Progress was recorded after each fixed block of 2,097,152 LCG iterations, allowing both wall-clock completion time and throughput to be measured. Thread CPU time was also recorded.

A separate Normal-priority periodic probe targeted 100µs intervals. Standalone background baselines were measured before and after the engine comparisons, with the same periodic probe present.

All paired generated **token IDs were identical**. Generated-token, MTP accepted/offered, and routing/cache summary counters also matched.

## Decode throughput

Values below are total generated tokens divided by total decode time within each condition.

| Continuous background threads* | N group | H group | Change |
|---|---:|---:|---:|
| 1 | 39.98 tok/s | 42.87 tok/s | +7.2% |
| 4 | 34.31 tok/s | 41.42 tok/s | +20.7% |

*Both conditions also included the periodic probe.

Increasing the background load from one to four busy threads reduced decode throughput by approximately **14.2% in the N group**, versus **3.4% in the H group**.

## Background-work throughput

Background throughput is normalized to the corresponding standalone baseline, using the average of the before/after baselines. Engine-run throughput is aggregated over the measured decode intervals.

| Continuous background threads | N group | H group |
|---|---:|---:|
| 1 | 98.0% of baseline | 97.3% of baseline |
| 4 | 98.2% of baseline | 94.3% of baseline |

Relative to N, H reduced busy-thread throughput by approximately **0.7% with one thread** and **4.1% with four threads**. Completing the same amount of work therefore took approximately **0.7% / 4.3% longer**, respectively, using the geometric mean of the paired ratios.

The throughput cost was modest in this workload, but responsiveness and throughput should be considered separately.

## Responsiveness and engine call tails

| Continuous background threads | Probe P99.9 lateness: N | Probe P99.9 lateness: H |
|---|---:|---:|
| 1 | 28µs | 193µs |
| 4 | 364µs | 1,959µs |

These are medians of the per-run P99.9 values, covering executed callbacks. Missed periodic slots are counted separately: with four busy threads, the median skipped-slot fraction increased from **1.28% to 17.86%**. Individual H runs also contained maximum probe delays of approximately **35–93ms**.

Conversely, the engine's CPU-pool call P99.9 decreased from approximately **13.56ms to 2.10ms** in the four-thread condition (medians of per-run call percentiles).

## Limitations

- One machine, one short-context workload, four independent process pairs per load condition, and 256 output tokens per run. This is preliminary evidence, not a universal performance or responsiveness guarantee.
- The 100µs probe remained continuously runnable using QPC and SwitchToThread, adding CPU contention; the same timer-arm overhead remained in all conditions. An initial sleeping-timer calibration did not achieve the required interval. This is **not a direct browser/UI responsiveness test**.
- Probe statistics cover an estimated decode interval reconstructed from the request-end timestamp and engine decode duration. CPU-call tail statistics cover all recorded multi-dispatch calls and are not timestamp-aligned to individual probe events.
- Priority changes affected both process class and thread-relative priorities; their individual contributions were not isolated.
- The pool's default 20ms spin-before-sleep policy remained unchanged. Its contribution to the observed trade-off has not been isolated. Bursty expert computation alone does not establish that idle gaps release the CPU to other programs.
- Identical MTP summary counters do not constitute a check of every individual draft sequence.

## Possible upstream direction

Would a **user-selectable Windows priority mode** be appropriate, retaining Normal priority as the default and preserving any existing relative thread-priority differences?

This is related to the CPU contention discussed in #921, but tests a different variable: **process/thread priority rather than spin duration**.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.