Pull requests / #1087

#1087 cpu: allow configuring expert-pool tasks per phase

closed · @Unmaple · 0 Kommentare · Auf GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsLinux

Beschreibung

Replaces #949, which GitHub closed after the upstream history rewrite. This branch starts at the new main (`82f46a8`) and contains only the contributor's task-count change: six files, +140/-8.

CPU expert row phases currently use three tasks per participating thread. Add `--pool-tasks N` so users can measure and select a grain for their workload. `0` retains the existing default; explicit values `1..4096` are capped by phase row count. Native sub-batches each use the target. The legacy single-token/oracle fallback is unchanged. The value is checked, documented, and printed at startup. This does not change CPU arithmetic, phase barriers, spin policy, GPU placement, or upstream `--host-core` behavior.

Origin and AI disclosure: the code changes were produced by **GPT-6-astra**, reviewed by the contributor, while using **Codex** to profile and tune Strata on an Intel Core i5-14600K (6 P cores + 8 E cores). Archived Gantt traces showed uneven task completion and idle tails. Finer partitions reduced those tails in that case; this observation does not establish a universal speedup.

Performance evidence (historical experiments; not rerun on the rewritten main): the nine-expert CPU-stage benchmark on v0.1.39 used 13 workers plus the host, seven quantization pairs and two routing patterns. 192 versus default 42 tasks reduced median layer-call time by 4.41%-21.00%. Each setting had 144 calls across three independent shuffled process runs; the process-level exact two-sided permutation p-value was 0.10. These are CPU-stage results, not a significant end-to-end gain. Full numerical tables and the task-tail chart are retained in #949. The chart illustrates P/E cooperation; execution restricted to P cores generally has a less pronounced completion gap.

No universal optimal task count or end-to-end speedup is claimed. The earlier combined-configuration observation involving worker count, PCIe policy, the soft-barrier approach related to #733, and IQ3_S P/E decode related to #863 is not an ablation of this option and is not used as evidence of its speedup.

Validation repeated on this PR head (`f7407d9`, based on rewritten upstream main `82f46a8`) on 2026-10-06:

- Clean-first full MSVC Release portable AVX2/CUDA sm_75 engine rebuild succeeded (143 build steps), using CMake's pinned llama.cpp and no local GGML override.
- Five targeted CTests passed: pool_tasks_test (168 synthetic bitwise comparisons), router_dot_parity, router_dot_parity_avx, pool_affinity_test, and pool_stress.
- Seven real quantization format combinations passed 1,176 bitwise partition comparisons, including host participation, repeated calls and native sub-batch boundaries.
- Ten pool-task CLI checks, 13 joint host-core/pool-task checks, and help-output checks passed using the newly rebuilt binaries. `git diff --check` passed.

Linux/AMD runtime behavior and the legacy non-native arithmetic path were not tested. The unchanged upstream expert_multi_test previously stopped at its AVX-512 capability guard on this AVX2-only host. These are build and correctness checks; no new throughput benchmark or full-model generation run was performed for this resubmission.

Related work: #500 moves intermediate activation quantization into a pool phase with an adaptive threshold; #733 makes native Down work available according to per-expert readiness; #863 concerns IQ3_S GU gathered decode on P cores and scalar decode on E cores. This PR only exposes row-task counts for the existing phases, retains the host quantization loop and phase barriers, and includes none of those changes. Integration with a future pipelined path needs a separate decision and validation. Automatic task tuning is outside this small change.

Mehr auf der Site

Links zu Install, Modellen, Releases.