Pull requests / #463

#463 decode: wait for an adaptive expert swap before reading the residency table

closed · @constantindjonkam · 0 comments · View on GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quants

Description

## What

Make greedy decode reproducible across requests, when an adaptive expert swap is in flight.
**Opt-in**, default off.

## The race

`apply_pending(false)` polls the swap's CUDA event with a **non-blocking** `cudaEventQuery`, so whether
the experts an adapt round had just copied in were *resident* for the next window depended on whether
their PCIe copies had landed.

A resident expert runs the GPU kernel; a streamed one is computed on the CPU. On a machine without
AVX-512 those two kernels round differently enough to move a top logit by **0.266** - from there the
trajectory is a different one, permanently.

Thank you for the independent measurement on 0.1.37 / RTX 5090 / NVFP4, which reproduces the race on a
different base, quant and tier. Same signature as here:

    same prompt, --expert-cache 8000, 400 tokens, 3 runs each
    without: 2 different outputs   with: 3 identical

## Why it is opt-in

Your numbers are the reason. The wait is bounded by the swap volume, so it is not free on a busy tier:
192 swaps every 2 rounds cost **~5% a round** (18.38 -> 19.32 ms). The same change here, on a lighter
tier (96 swaps every 4 rounds, 23-77 per round), was **under 0.3%** (78.27/78.19 -> 78.30/78.06 tok/s).

So it is a real trade, and the right answer is not one size for all: `STRATA_ADAPT_WAIT=1` asks for
reproducibility, and an install's default speed is unchanged.

Worth keeping in view: the wait is not only a cost. It also makes the swaps *count* sooner - your hit
rate 0.898 -> 0.906 and misses/layer 3.11 -> 2.92, with the CPU pool at 6.5-8.0 ms against 7.3-8.1.
On a tier that busy the 5% buys fewer CPU experts.

This matches the flag you already carry in your fork, so the two do not have to be reconciled.

## Scope

`apply_pending(adapt_wait)` in both decode loops. Nothing else changes: no new counters, no change to
the default path, and the residency table is untouched.

The `STRATA_TRACE_ADAPT` lines that show the race are env-gated and off by default; they ride along
here so the PR is self-verifying.

## Reproducing

    STRATA_ADAPT_WAIT=1   # identical requests then repeat; without it, a busy tier will fork

The per-window probe in [#462](https://github.com/Niko1221/Strata/pull/462) localises it further if
you want the window index rather than just the divergence.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.