Pull requests / #501
#501 verify: pin the serve host thread the session loop already pins
closed · @yannickloth · 0 评论 · 在 GitHub 查看
BenchmarksServer & APIMulti-GPU
描述
Pin the serve/spec host thread the way the session loop already pins its host. ## What this targets `SessionLoopScratch::init` pins the host thread to the core `ExpertPool` reserves (`physical_cores(false)[0]`; workers come from `physical_cores(true)`), because the host joins the pool drain with the default `host_works` (`generate.cpp:2637`) - the startup log's "8 expert-pool workers + the host thread". `Verifier::run` (`verify.cpp`) spins on the doorbell and then calls the same pool, but never pinned, so its host was free to land on a worker's core or SMT sibling. See #494: the CPU pool is 53% of a fresh round / 26% warm, and the host is one of the ~9 draining threads. ## Measured first (#494): no collision on this host On production 0.1.35 there is no collision to fix: `SessionLoopScratch::init` is called unconditionally in `main()` (`generate.cpp:3607`) *before* the serve/verifier block (`:4200`), so the serve host is already on the reserved core. During an active serve decode the only busy non-worker thread is the engine main thread, with `Cpus_allowed_list: 0`, and the 8 workers on `2,4,6,8,10,12,14,16`; 120 x 50 ms samples give exactly nine busy threads, the host on core 0 in every sample. ## The fix `Verifier` now carries the same pin/restore as `SessionLoopScratch`: - `host_pin_core(split)`: the first physical core `physical_cores(true)` drops (never a worker's), or -1 under a layer split / when no such core exists. - `maybe_pin_host()` once at the top of `run()`, before the spin: `STRATA_NO_VERIFY_PIN` is the off arm; refuses when the calling thread's affinity does not include the core; pins once and records the previous mask and thread id. - the destructor restores the affinity only when the same thread destroys the verifier. - a layer split never pins: all stages would fight for the one reserved core and serialize. No numerics change, no token-loop change, no pool worker-count or affinity change. ## A/B (same binary, `STRATA_NO_VERIFY_PIN`) Local `build/strata`, `bench/e2e.sh`, fixed prompt, `max_tokens 200`, interleaved, 6 runs per arm, alternating order: | arm | runs (tok/s) | median | mean | stdev | | --- | --- | ---: | ---: | ---: | | off | 31.5, 22.7, 22.7, 26.9, 25.8, 25.7 | 25.75 | 25.88 | 2.97 | | on | 32.8, 34.7, 18.8, 19.7, 24.5, 24.4 | 24.45 | 25.82 | 6.03 | Mean delta **-0.3%**, median -5.1%: *not* the >3% win `tools/calibrate.py` requires (the host is noisy, +/-20%). This is a **robustness fix, ~0 perf**: it makes the reservation local to the verifier instead of depending on the session scratch that happens to run first. ## Evidence - Placement (pin on, local build, serve): the host thread is on core 0 in every sample and the workers on 2..16; `/proc/<pid>/task/*/status` confirms. - Fixed-prompt output **bit-identical** with the switch on vs off (306 chars). - `verify_host_pin_test`, `pool_test`, `pool_stress`, `expert_parity`, `file_expert_source_test`, `expert_layout_test`, `expert_cache_per_layer_test` and 10 CPU/GPU parity tests pass. `expert_multi_test` fails its **pre-existing** no-AVX-512 refusal (this host has AVX2 + AVX-VNNI). - `verify_host_pin_test` asserts the chosen core is not a worker core, and that a layer split returns -1. ## Caveat This is a single-GPU host, so the layer-split guard is verified by inspection and by the unit test, not by running a multi-GPU split. `maybe_pin_host` skips whenever `next_ != nullptr || lb_ > 0 || le_ < g.n_layers`, i.e. every stage. Refs #494, #489, #485.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。