Pull requests / #905

#905 Pipelined windows + async adaptive tier together (stacked on #859 and #876)

closed · @Hardin22 · 0 commentaires · Sur GitHub

BenchmarksSetup & installModels & quantsWindows

Description

**Stacked on #859 and #876.** The branch carries their commits; only the last commit (`887a711`) is new. I'll rebase it once those land.

With the pipelined windows (#859), two windows can be in flight on a two-card split. The async adaptive tier (#876) moves experts between RAM and the cards in steps that advance between windows. Each feature alone keeps the other out. This commit lets them run together, with fences:

- **Ticks**: inside the pipelined loop, the async tier ticks once per verdict.
- **Fences**: a fence is an event recorded on both stages' streams. Both parities' verifiers share their stage's stream, so one fence covers every window and commit queued on both cards at that moment.
- **Step 2** (copy the new experts into the evicted slots): runs only after the windows that still saw the old residency have passed the fence (`fence_ev`).
- **Step 3** (move the evicted blobs into the RAM places): waits for a second fence (`exch_ev`) before the RAM copy is committed.
- **CPU side**: the pool runs synchronously inside `service()`, so it needs no fence. The pipelined verifiers keep the always-copy doorbell publish.
- **End of a pipelined loop** (after `pl_drain`): the job is waited for and the fences dropped. The next request's drain finishes the round. An error that ends the loop makes a job waiting for a fence give up.
- **Fallback**: `STRATA_PIPELINE_ADAPT_ASYNC=0` keeps the blocking tier beside the pipelined windows (the startup log says so).

## Measured

4060 Ti (layers 0-19) + 5080 (20-47) split, i9-14900KF, 32 GB of RAM, resident RAM mode (#848). Swift 1.5 IQ3_XXS at 160K, with a fine-tuned draft layer, `--spec-min-p 0.65`, `--trim-stage-weights`, and Windows power throttling off. Greedy, five prompts × two rounds, decode tok/s:

| | runs |
|---|---|
| `--pipeline-windows 2` with the blocking tier | 111.2, 111.4 |
| `--adapt-async 1`, serial loop | 117.3, 109.6 |
| **both (this PR)** | **121.7, 126.3, 120.9** |

That's about +10% over the pipeline alone and +8% over the async tier alone. One more run of the combination came in at 105.4: its first prompt was cold, at 64 tok/s.

## Checks

- **Stress, both on**: no degenerate answer and no stall in either run.
  - 22 varied requests at temperature 0.6 (thinking on/off, five languages, code), with 48,387 experts swapped in over 658 rounds.
  - 11 more with `STRATA_PIPELINE_FORCE_MISS=2 STRATA_PIPELINE_THETA=0` (every second guess wrong, every window speculated).
- **No exact check**: the async tier is not bit-exact by design (#876), so there is no exact comparison.

For scale, on this PC: setup's choice today (the 5080 alone) decodes ~29 tok/s. With #848 + #859 + #876 + this + #904, the same split reaches ~125.

Sur le site

Liens install, modèles, releases.