Pull requests / #501

#501 verify: pin the serve host thread the session loop already pins

closed · @yannickloth · 0 comentários · No GitHub

BenchmarksServer & APIMulti-GPU

Descrição

Pin the serve/spec host thread the way the session loop already pins its host.

## What this targets

`SessionLoopScratch::init` pins the host thread to the core `ExpertPool` reserves
(`physical_cores(false)[0]`; workers come from `physical_cores(true)`), because the host
joins the pool drain with the default `host_works` (`generate.cpp:2637`) - the startup
log's "8 expert-pool workers + the host thread". `Verifier::run` (`verify.cpp`) spins on
the doorbell and then calls the same pool, but never pinned, so its host was free to land
on a worker's core or SMT sibling. See #494: the CPU pool is 53% of a fresh round / 26%
warm, and the host is one of the ~9 draining threads.

## Measured first (#494): no collision on this host

On production 0.1.35 there is no collision to fix: `SessionLoopScratch::init` is called
unconditionally in `main()` (`generate.cpp:3607`) *before* the serve/verifier block
(`:4200`), so the serve host is already on the reserved core. During an active serve
decode the only busy non-worker thread is the engine main thread, with
`Cpus_allowed_list: 0`, and the 8 workers on `2,4,6,8,10,12,14,16`; 120 x 50 ms samples
give exactly nine busy threads, the host on core 0 in every sample.

## The fix

`Verifier` now carries the same pin/restore as `SessionLoopScratch`:

- `host_pin_core(split)`: the first physical core `physical_cores(true)` drops (never a
  worker's), or -1 under a layer split / when no such core exists.
- `maybe_pin_host()` once at the top of `run()`, before the spin: `STRATA_NO_VERIFY_PIN`
  is the off arm; refuses when the calling thread's affinity does not include the core;
  pins once and records the previous mask and thread id.
- the destructor restores the affinity only when the same thread destroys the verifier.
- a layer split never pins: all stages would fight for the one reserved core and
  serialize.

No numerics change, no token-loop change, no pool worker-count or affinity change.

## A/B (same binary, `STRATA_NO_VERIFY_PIN`)

Local `build/strata`, `bench/e2e.sh`, fixed prompt, `max_tokens 200`, interleaved, 6
runs per arm, alternating order:

| arm | runs (tok/s) | median | mean | stdev |
| --- | --- | ---: | ---: | ---: |
| off | 31.5, 22.7, 22.7, 26.9, 25.8, 25.7 | 25.75 | 25.88 | 2.97 |
| on | 32.8, 34.7, 18.8, 19.7, 24.5, 24.4 | 24.45 | 25.82 | 6.03 |

Mean delta **-0.3%**, median -5.1%: *not* the >3% win `tools/calibrate.py` requires
(the host is noisy, +/-20%). This is a **robustness fix, ~0 perf**: it makes the
reservation local to the verifier instead of depending on the session scratch that
happens to run first.

## Evidence

- Placement (pin on, local build, serve): the host thread is on core 0 in every sample
  and the workers on 2..16; `/proc/<pid>/task/*/status` confirms.
- Fixed-prompt output **bit-identical** with the switch on vs off (306 chars).
- `verify_host_pin_test`, `pool_test`, `pool_stress`, `expert_parity`,
  `file_expert_source_test`, `expert_layout_test`, `expert_cache_per_layer_test` and 10
  CPU/GPU parity tests pass. `expert_multi_test` fails its **pre-existing** no-AVX-512
  refusal (this host has AVX2 + AVX-VNNI).
- `verify_host_pin_test` asserts the chosen core is not a worker core, and that a layer
  split returns -1.

## Caveat

This is a single-GPU host, so the layer-split guard is verified by inspection and by the
unit test, not by running a multi-GPU split. `maybe_pin_host` skips whenever
`next_ != nullptr || lb_ > 0 || le_ < g.n_layers`, i.e. every stage.

Refs #494, #489, #485.

No site

Links install, modelos, releases.