Pull requests / #904

#904 Verify window: PDL on sm_90+, graph branches, batched Q4_0 KV append and PLE post-ops (bit-identical)

closed · @Hardin22 · 0 comentarios · En GitHub

BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Descripción

This is the part of the fork's "fewer, fuller launches" work that main doesn't have yet. It comes as two commits, and each piece has its own switch. Every piece keeps the arithmetic of the path it replaces, so the output is unchanged; the checks are below.

1. **Programmatic dependent launch on sm_90+** (`STRATA_DF_PDL`). A decode window is a long chain of small kernels, and on a fast card many of them finish before the next one has fetched its weights. With PDL, the quantize kernel, the multi-column MMVQ and the multi-row BF16 GEMV load their weights before `griddepcontrol.wait`.
   - **Where it applies**: the attribute is only set while capturing, and only when every predecessor of the next node is a kernel (never after a memcpy, memset, host or event node). `STRATA_DF_PDL=2` restricts it to single-predecessor edges.
   - **When it is on**: `pdl_supported()` is false on HIP and on CUDART < 12.3. It also needs `cc_major >= 9` and a binary built for sm_90+ (PTX JIT from older code would lack the wait).
   - **Other changes**:
     - `copy_rows_strided` replaces a memcpy2D node.
     - The three PDL kernels drop `__restrict__` (llama.cpp's PDL fix), which can change speed but not values.
2. **Graph branches and batched forms**:
   - `STRATA_DF_BRANCH`: the GDN a/b/z, query and indexer-query sides and the step/pos inputs on a side branch. **Opt-in** since f89d31f: on Linux (open modules, GSP) it stalled with NVML queries between requests (#905).
   - `STRATA_DF_QB_Q4`, `STRATA_DF_KVAPP`: the Q4_0 KV's window batched (main still appends per token with `--kv q4_0`).
   - `STRATA_DF_PLE`: the PLE post-ops in one pass.

Already in main, so left out: `native_swiglu_quantize_q8_1` / `sigmoid_scale_rows`, fd95405's attention change, the shared-expert branch, and STRATA_PLE_BATCH's projections. I also dropped a one-row-per-warp MMVQ variant, because it did not beat main's ROWS=2.

## Measured

Swift 1.5 IQ3_XXS at 160K (q4_0 KV), RTX 4060 Ti + RTX 5080 layer split, with the resident RAM mode (#848), pipelined windows and the async tier (#859, #876). Greedy, five prompts × two rounds, decode tok/s.

| | run 1 | run 2 |
|---|---|---|
| every switch on (default) | 123.5 | 119.7 |
| every switch off (`STRATA_DF_PDL=0 STRATA_DF_BRANCH=0 STRATA_DF_QB_Q4=0 STRATA_DF_KVAPP=0 STRATA_DF_PLE=0`) | 118.1 | 117.5 |

About **+3%**. PDL only acts on the 5080 (sm_120); the 4060 Ti (sm_89) has none.

## Checks

- **Exact pass** (`STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`, serial loop, 4 prompts × 500 tokens): all switches on and all off give the **same text**.
  - Against a build without these commits, 2 of 4 differ. The larger binary leaves 2 fewer expert slots on each card (5740 vs 5742, 4762 vs 4764), so a few experts move to the CPU, which rounds them differently.
- **Unit tests** pass on both cards (RTX 5080, RTX 4060 Ti):
  - `pdl_parity`: 0 of 200 replays differ; programmatic edges only go kernel → kernel; also with `STRATA_DF_PDL=2`.
  - New: `verify_batch_parity` (every batched form against its per-token path, bitwise).
  - Existing: `mmvq_multi_parity`, `kv_q4_parity`, `bf16_gemv_parity`.

## Not tested

- HIP and CUDA 12.3–12.9 builds were not compiled here.
  - HIP takes the plain launch.
  - `griddepcontrol` is only compiled under `!__HIPCC__ && __CUDA_ARCH__ >= 900`.

En el sitio

Enlaces a install, modelos, releases.