Pull requests / #733
#733 Pipeline CPU expert work without a batch-wide phase barrier
closed · draft · @moorwu · 0 评论 · 在 GitHub 查看
BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux
描述
# Pipeline CPU expert work without a batch-wide phase barrier ## Why this helps CPU expert verification currently finishes Gate/Up for the whole batch, quantizes intermediate activations, then starts Down. Experts that are ready early must wait for the rest of the batch. This opt-in change quantizes each expert as soon as its last Gate/Up tile finishes and publishes that expert's Down work. All producer tiles are claimed before dependent tiles, so dependent waits cannot strand unclaimed producers. The existing row partitions and arithmetic kernels are retained. ## Use Set `STRATA_POOL_PIPELINE=1` before constructing the expert pool. It is off by default and applies to the native multi-token path. ## Validation environment - Two modified RTX 3080 20 GB cards; CUDA SM86, CUDA 13.1, Linux. - Intel Xeon E5-2686 v4 (18 cores / 36 threads), about 94 GiB usable RAM. - PCIe 3.0, no GPU peer-to-peer access; 17 expert-pool workers. - Qwen3.8-Flash-Next IQ3_S, 131,072-token context, INT8 KV, 32,768 resident KV cells. - Medium thinking: 2,048-token reasoning cap, 8,192-token output cap. - Physical GPU power limits stayed at 250 W and 280 W. No power increase. - Historical performance tests used an isolated engine based on official v0.1.38 plus reviewed architectds/Strata changes (best, 05c0f36). This PR contains only our change, ported to official main (99f3dbd). The historical timings are NOT a measured speedup of this pure-upstream port. ## Historical benefit This is a small scheduling improvement, not a large tokens/s claim: | With disjoint cache and helper 9500 | Pipeline off | Pipeline on | | --- | ---: | ---: | | 32K complete structured request | 11.128 s | 11.020 s | | Natural-language generation | 76.5 tokens/s | 77.2 tokens/s | The extra gain was about 1%. The natural-language outputs differed in length; their wall-time difference must not be described as equal-work acceleration. Both variants passed complete-answer and tool-call checks. Broader cache gains are not attributed to this PR. ## Validation The historical real-model test passed 810/810 bitwise comparisons across three layer formats, 1/4/17 workers, host participation on/off, 1/2/4 tokens, 1/2/7/18/97 jobs, and three repeats. An independent serial row reference also passed. This is not a ThreadSanitizer or general model-quality result. The included real-model test can be run as: ```sh pool_pipeline_parity /path/to/native/pack /path/to/corresponding/model-shard.gguf ``` Without model arguments it exits 77 (CTest skip), not a false pass. It reads the first three layers from the supplied shard; use the matching Flash-Next pack and shard. The independent pure-upstream port also passed the 810/810 real-model bitwise comparisons on the stated Linux host. All known expert-pool header consumers were recompiled, and the independent engine and parity test linked successfully. The older upstream VNNI pool selftest was skipped because this CPU lacks the required AVX-512 features; it was not counted as passing. No model server was started or restarted, and no production settings were changed. ## Related work and remaining work PR #500 parallelizes intermediate quantization as a separate phase. This proposal instead pipelines per-expert readiness across phases. They touch the same code and should be coordinated rather than merged blindly. Keep draft until the independent upstream port completes end-to-end performance and concurrency validation. No ThreadSanitizer run has been performed. Other CPU architectures are unvalidated. ## 0.1.39 follow-up The branch now includes official 0.1.39 (`6f32ec0`). Its refactored row and quantization helpers retain upstream `q2_native_kernels(...)` ISA dispatch instead of choosing kernels from the format number alone. The parity reference uses the same dispatch. The test is registered on Windows and Unix, with Windows environment-variable and 64-bit file-offset equivalents. A contributor reported 810/810 comparisons passing on Ryzen 9 7940HS with Windows/MSVC, but no convincing speedup: pipeline-on decode was 31.8 tokens/s against off controls of 33.4 and 30.6. That was a combined build, not this exact updated branch. The report is linked in PR #733; credit belongs to @midhatn. Our combined custom 0.1.39 build passed a five-arm, 240-request HTTP suite. Its whole-build gains cannot be attributed to this small scheduling change. Neither that suite nor bitwise parity is a ThreadSanitizer result. Keep the PR draft until isolated performance and concurrency validation is convincing. The exact revised branch compiled and linked independently against official 0.1.39 on our Linux host. All known expert-pool header consumers were rebuilt. Using the matching SC117 abliterated IQ3_S pack and shard, the updated parity test passed 810/810 comparisons, including the independent serial reference, 1/4/17 workers, host participation on/off, repeated calls and 97-job batches. We did not run the revised branch on Windows or launch a model server for this check. The contributor's Windows result is separate evidence, not our own run.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。