Pull requests / #859

#859 Layer split: pipelined verify windows for one conversation (--pipeline-windows, opt-in)

closed · @Hardin22 · 0 comentarios · En GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsWindows

Descripción

Refs #642. These are the pipelined windows from the fork, ported onto main and **opt-in** (`--pipeline-windows 2`; the default 0 changes nothing).

With a layer split, one conversation's decode is serial: stage 0 runs its layers and hands off, then waits while the last card runs its layers, the head and the draft. The batch pipeline from #559 overlaps *different* conversations, so a single conversation gets nothing from it.

`--pipeline-windows 2` starts window K+1 on stage 0 while stage 1 still verifies window K.
- **The guess**: K+1 is guessed by running the MTP chain through K's drafts (teacher-forced), launched as soon as K's verdict is known.
- **Guess right**: half of K+1 is already done.
- **Guess wrong**: stage 0 restores its GDN state from a snapshot, replays K's commit with the real count, and K+1 is rebuilt from the real tokens.
- **When it speculates**: only when the chain's estimate says K+1 is likely to be kept (`STRATA_PIPELINE_THETA`).

Two commits:
1. **MtpDrafter**: a teacher-forced chain as one asynchronous launch. Nothing calls it on its own, and `draft()` is unchanged for the serial loop.
2. **The pipelined loop**: a non-blocking form of the Verifier's `run()`/`commit()` on the `batch_launch`/`batch_poll` pattern, two verifiers per stage, the adaptive tier fenced against the windows in flight, and the VRAM for the second verifier and the snapshots. It turns itself off, with one log line, for anything it does not handle:
   - `--batch`;
   - `--peer-device`;
   - helper caches / `--remote-expert-opt`;
   - three or more stages, or a split on one GPU;
   - no drafter.

   Requests with `penalty_last_n` or coupled draft sampling decode serially.

## Measured

Swift 1.5 IQ3_XXS at 160K (q4_0 KV, stock draft layer). RTX 4060 Ti (layers 0-19) + RTX 5080 (20-47), i9-14900KF, 32 GB of RAM, with the resident RAM mode on the split from #848. Greedy, 500 tokens, two interleaved serial/pipelined pairs of three rounds each, decode tok/s:

| | serial | `--pipeline-windows 2` | |
|---|---|---|---|
| Python code | 90.4 | 103.1 | +14% |
| C code | 76.6 | 81.6 | +7% |
| English prose | 63.0 | 71.1 | +13% |
| Italian prose | 41.6 | 46.3 | +11% |
| copy-heavy edit (a 5 KB file back with one rename) | 90.5 | 117.9 | +30% |
| **mean** | **72.4** | **84.0** | **+16%** |

It pays when the windows are GPU-bound. On the same PC without #848 (experts read through the OS file cache, `--mmap-experts`), the file reads dominate and it was not measurably faster.

## Correctness

- With `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`, the greedy text is **identical bit for bit** to the serial loop's, as long as both have the same expert caches: 4/4 prompts, 500 tokens. The serial run was given the pipeline's reserve through `--vram-reserve-mib`.
- The same text also comes out with each of these:
  - `STRATA_PIPELINE_THETA=2` (no speculative window at all);
  - `STRATA_PIPELINE_FORCE_MISS=1` (every speculative window wrong and rolled back), with and without `STRATA_PIPELINE_DOOM_SKIP=0`;
  - `STRATA_PIPELINE_PRESTAGE=0`;
  - `STRATA_PIPELINE_SNAP_OVERLAP=0`.
- Against a run *without* the flag, the pipeline's 253 MiB (first card) + 160 MiB (last card) move a few experts from a card to the CPU, which rounds them differently, so a near-tie can flip. All the flips I looked at are word-level alternatives ("components" vs "year, month, day").
- With the adaptive tier on, no degenerate answer and no stall in either run:
  - 22 varied requests at temperature 0.6 (thinking on/off, five languages, code);
  - 11 more with `STRATA_PIPELINE_FORCE_MISS=2`.

## Not tested

- **HIP and SYCL**: not built here.
  - On HIP, the only new CUDA-specific call (the drafter's stream priority) is behind `#if !defined(STRATA_USE_HIP)`, and the new forcing kernel is plain code.
  - SYCL's migrated verify/mtp would need the new methods before it can use the flag. It doesn't build them today.
- **The zero-doorbell graph** (every expert resident in VRAM): not exercised on this PC. There, `service()` only raises the stage flag.

Happy to split it differently or change the defaults if you prefer.

En el sitio

Enlaces a install, modelos, releases.