Pull requests / #949
#949 cpu: allow configuring expert-pool tasks per phase
closed · @Unmaple · 0 comentários · No GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux
Descrição
CPU expert row phases currently use three tasks per participating thread. That fixed grain can leave useful CPU-layer latency reductions unavailable to users tuning their worker count. Origin and AI disclosure: the code changes were produced by **GPT-6-astra** while the contributor used **Codex** to test and tune Strata on an Intel Core i5-14600K (6 P cores + 8 E cores). CPU expert processing did not reach the expected aggregate throughput. Task-timeline/Gantt charts showed uneven work completion and idle tails across threads. Scheduling overhead appeared small compared with idle time, suggesting that the default task count was too coarse for this workload. Increasing the number of row partitions produced repeatable CPU-stage improvements on this host. This is the working explanation from the profiling experiments; the measurements do not isolate task count as the sole cause or establish a benefit on every hybrid CPU.  Earlier instrumented experiment: 6 P cores + 1 E core, nine experts, 14 token routes. Increasing tasks per phase from 21 to 84 reduced the first-to-last lane completion gap from 180.3 to 39.1 μs for Gate/Up, and from 86.2 to 11.3 μs for Down. These are two individual archived traces, separate from the 14-participant benchmark below.This illustrates cooperation between P and E cores; when execution is restricted to P cores, the completion gap is generally less pronounced. Add `--pool-tasks N` for the batched CPU expert paths. The default `0` preserves the existing policy. Explicit values `1..4096` are capped by each phase's row count; native sub-batches each use the target. The legacy single-token/oracle fallback is unchanged. The setting is documented and reported at startup. Automatic tuning, phase pipelining, kernels, and GPU placement are outside this change. On an i5-14600K with 13 workers plus the host, nine real experts, seven quantization combinations, and two token-routing patterns, the preselected 192-task setting reduced median CPU-layer call time by 4.41%-21.00% versus the default 42 tasks. Each configuration has 144 samples across three shuffled process runs. 192 was not always the best of 128/192/256. These are CPU-stage measurements; no end-to-end tokens/s gain attributable to this patch alone, or cross-machine gain, is claimed. The CPU-stage measurements are listed below. An isolated end-to-end comparison for this patch remains pending; a separate combined-configuration experiment is described below. Validation: - Windows/MSVC Release portable AVX2 build of the full CUDA engine succeeded. - New CPU-only synthetic Q4_0 CTest: 168 bitwise partition comparisons passed. - Seven real format combinations: 1,176 bitwise comparisons passed, including mixed token counts, host participation off/on, repeated calls, and crossing the native sub-batch boundary. - Invalid CLI values and constructor bounds checked; `git diff --check` passed. Linux/AMD runtime behavior and the legacy non-native arithmetic path were not tested here. Existing default scheduling is retained. Related work: [#500](https://github.com/Niko1221/Strata/pull/500) dispatches intermediate activation quantization as a separate pool phase, controlled by an adaptive threshold (`STRATA_POOL_QUANT_THRESH`). [#733](https://github.com/Niko1221/Strata/pull/733) pipelines native work according to per-expert readiness, allowing an expert's Down work to become available before the whole batch finishes Gate/Up. This patch only exposes the row-task count for the existing Gate/Up and Down phases; it retains the host quantization loop and existing phase barriers and includes neither proposal. All three touch the CPU pool, including `run_split_multi` and/or `run_split_multi_native`. Nearby edits may require conflict resolution depending on merge order; this has not been tested as a combined patch. If either scheduling proposal lands first, the task-count changes should be rebased onto its final structure, with correctness and performance checks repeated. The `--pool-tasks` value governs row tasks, not #500's expert-token quantization jobs. With #733, whether and how the option governs the pipelined path needs an explicit integration decision. The measured gains here are against upstream v0.1.39 without either proposal; no combined speedup is claimed. Measured median CPU-layer call time (microseconds): | GU / DOWN quantization | Routing | Default (42) | 128 | 192 | 256 | |---|---|---:|---:|---:|---:| | IQ3_XXS / IQ4_NL | p0 | 512.35 | 490.80 | 485.75 | 479.65 | | IQ3_XXS / IQ4_NL | p1 | 523.10 | 501.05 | 499.20 | 498.50 | | IQ2_S / Q2_0 | p0 | 372.50 | 357.40 | 352.60 | 343.10 | | IQ2_S / Q2_0 | p1 | 423.20 | 402.70 | 401.70 | 394.10 | | IQ2_S / IQ4_NL | p0 | 467.50 | 450.50 | 438.60 | 430.95 | | IQ2_S / IQ4_NL | p1 | 513.90 | 481.00 | 484.05 | 474.90 | | IQ3_XXS / Q2_0 | p0 | 416.00 | 401.00 | 393.75 | 390.65 | | IQ3_XXS / Q2_0 | p1 | 433.55 | 424.45 | 414.45 | 414.15 | | IQ3_S / IQ4_NL | p0 | 705.20 | 582.90 | 569.00 | 582.90 | | IQ3_S / IQ4_NL | p1 | 643.40 | 602.65 | 591.50 | 581.80 | | IQ3_S / Q2_0 | p0 | 609.10 | 494.70 | 481.20 | 488.30 | | IQ3_S / Q2_0 | p1 | 548.10 | 523.60 | 510.10 | 501.45 | | IQ4_XS / IQ4_NL | p0 | 627.20 | 589.55 | 586.50 | 578.40 | | IQ4_XS / IQ4_NL | p1 | 633.00 | 577.40 | 574.05 | 570.35 | p0 uses one input token with nine expert routes. p1 uses five two-token experts and four one-token experts (14 routes). The 13 workers and caller were pinned. Each independent process used eight warm-up calls and 48 measured calls; three runs per configuration were shuffled with seed 20261005. Weight cache lines were flushed before each measured call, outside the timer. The timer includes activation quantization, job setup, and CPU pool execution; routing, PCIe transfers, GPU work, and final merging are excluded. Outputs were checked bitwise against the default partitioning. Statistical limitation: there are only three independent process runs per configuration. Although all three 192-task process medians were below all three default process medians in each case, the exact two-sided permutation p-value at the process level is 0.10. This is repeatable evidence on this host, not a claim of significance at the 5% level; the 144 calls are not treated as 144 independent process samples. Independent clean-build check: a fresh clone of the committed patch and a new build directory successfully built the full CUDA engine using CMake's automatically fetched llama.cpp pin (3cf03257f219afbe7334045ff7c6a06ac68c627d), without a local GGML override or external benchmark target. Engine help, 10 CLI checks, the 168 synthetic partition comparisons, and 1,176 real-format partition comparisons passed with these fresh binaries. The existing router parity (including AVX), affinity, and pool stress tests also passed. One additionally selected upstream test, expert_multi_test, exited at its unchanged AVX-512 capability guard on this AVX2-only host; its test source and expert.cpp are unchanged by this patch. This does not claim all upstream tests pass on this machine. Full model generation remains pending. Separate combined-configuration observation (not an ablation of this PR): a related experimental engine using 14 participants, 192 nominal row tasks, per-expert pipelining (the soft-barrier approach related to #733), and IQ3_S GU gathered decode on P cores / scalar decode on E cores (related to #863), with PCIe fraction 0.14, was compared with the production v0.1.38 engine using six participants, 18 default tasks, and PCIe fraction 0.36. Across seven alternating AB/BA process pairs and 42 measured requests cycling through mathematics, Python, and SQL, aggregate decode throughput was 42.05 vs 44.30 tokens/s (+5.33%; paired group bootstrap 95% interval +4.43% to +6.21%, exact two-sided sign-flip p=0.0156 under its exchangeability assumption). This measures the entire configuration bundle, including the CPU-worker and PCIe-policy changes; it does not isolate task granularity, pipelining, or gather, and has not established that the tested binary is exactly the three published PR heads combined. Gains depended on the workload/cache protocol (the warmed first-topic subgroup was +1.32%). There was no fixed-output replay, and alternating order does not eliminate all system drift. These observations motivate a controlled ablation; they are not evidence of a +5.33% gain from this patch alone.
No site
Links install, modelos, releases.