Pull requests / #1122

#1122 Pipelined windows + async adaptive tier together (--pipeline-windows 2 --adapt-async 1)

closed · @Hardin22 · 0 Kommentare · Auf GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindowsLinux

Beschreibung

This replaces #905, which GitHub closed when `main` was force-pushed. #859 and #876 are in 0.1.40, so it is now one commit on the new `main` (0.1.40.1).

0.1.40 turns `--adapt-async` off when `--pipeline-windows 2` is set. This commit lets them run together. On a 4060 Ti + 5080 split that is +18% over the pipeline alone and +11% over the async tier alone.

## How

With `--pipeline-windows 2` there is always a window in flight when the tier ticks (window K+1 on stage 0 while stage 1 verifies or commits K). Each window waits on the loop's thread at every layer and plans each layer from the residency table as it is when that layer is served. Two rules keep the tier out of its way:

- **No CUDA call from the job thread while a window waits.** A copy queued to a card whose window spins on a host flag can block inside the driver until that window ends (issue #31), and a blocked thread can hold up the loop the window is waiting for. So beside the windows the job thread does host work only (the choice, the bounce copies, the RAM moves). The loop's thread queues each card's copies itself at that card's next gap: a moment its stage has no window in flight, before a window is launched there and at every tick.
- **A gap after a step is that step's fence for the card.** Every window planned there before the step has completed by then.
  - Step 2's copies into the evicted slots go at a gap after step 2, since a window planned before may still read the old expert from its slot.
  - Step 3's RAM moves (a new `RamFence` state) wait for a gap after step 3 on every card whose layers exchanged, since a window planned before may still pull an incoming expert over PCIe from the RAM place its evicted one takes. With `STRATA_EXCHANGE_ROTATE`, the flip that hands that place over comes after them too.
  - No events and no host waits. The CPU's reads need no fence, because the pool runs inside the loop's service calls.

The tier's steps advance at every verdict, after the drafter's chain is launched, and every 0.5 ms between verdicts while a round is in flight. The device copies of the residency table are uploaded once, when the loop has drained (the pipelined verifiers publish every row anyway). At the loop's end, whatever still waits for a gap is queued and the round in flight finishes at the next request's drain.

`--pipeline-windows 1` and exchange buffers that are not page-locked still turn the async tier off. `STRATA_PIPELINE_ADAPT_ASYNC=0` (read under `STRATA_PIPELINE_DEBUG=1`, like the other pipeline test variables) keeps the blocking tier beside the windows. The serial loop's async tier is unchanged.

Also, as `run()` already does: on a timeout the pipelined `Verifier::service` prints the verifier's diag before it releases the GPU waits, and the trace after it with `STRATA_VERIFY_TRACE=1`.

## Measured

0.1.40, RTX 4060 Ti 16 GB (layers 0-19) + RTX 5080 16 GB (20-47), i9-14900KF, 32 GB of RAM, Windows 11. Resident RAM mode with the whole complement in RAM, Swift 1.5 IQ3_XXS at 160K (q4_0 KV), stock draft layer, `--trim-stage-weights`. Greedy, five prompts × two rounds, a fresh server per run, decode tok/s:

| | run 1 | run 2 |
|---|---:|---:|
| `--pipeline-windows 2` (0.1.40: blocking tier) | 93.5 | 94.1 |
| `--adapt-async 1`, serial loop (0.1.40) | 97.4 | 102.1 |
| **both (this PR)** | **110.2** | **111.3** |

## Checks

- **Stress, both on**, with `STRATA_PIPELINE_FORCE_MISS=2 STRATA_PIPELINE_THETA=0` (every second guess wrong, every window speculated): 12 sequential 1500-token coding requests at temperature 0.6, with an `nvidia-smi` query after each, as in the stall report on #905. No timeout and no degenerate answer. The tier ran 1552 rounds and swapped in 105,435 experts at about 40 ms per round. (That stall was later traced to #904's graph branches plus NVML queries on Linux; `STRATA_DF_BRANCH` is opt-in there now.)
- **No exact check.** The async tier is not bit-exact by design (#876). On 0.1.40 the serial loop also gives a different greedy text for a different window size (`--spec 4` vs `--spec 2`: 2 of 4 prompts identical), so pipelined vs serial can't be compared bit for bit either.
- Builds with MSVC + CUDA 13.3 (sm_89 / sm_120).

The bounded attention merge and the pipelined PLE switches (#910) are stacked on this and will follow as their own PR.

Mehr auf der Site

Links zu Install, Modellen, Releases.