Pull requests / #302

#302 HIP/Windows: batch expert copy dependencies in long prefill

closed · @araujoluks · 0 comments · View on GitHub

BenchmarksServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

Description

## What

Long-prefill HIP/Windows issued completion/reuse dependencies at expert granularity. This groups those dependencies while preserving the individual expert copies, sources, destinations, bytes and logical order.

## Why

On one RX 9070 XT / gfx1201 / HIP 7.2 / Windows, dependency management substantially reduced effective prefill throughput. This changes dependency granularity; it does not combine transfers or change kernel arithmetic.

## Implementation

- Direct pinned groups use one terminal `copied` record and one terminal `used` reuse wait.
- Host staging groups use one terminal `dma_done` record. Each reused host buffer still waits for DMA completion before CPU memcpy.
- Pinned staged GPU copies independently group their `copied`/slot dependencies.
- Groups stop at incompatible GPU/host wraps and source-type boundaries. Publication follows the terminal record. Prior compute waits/uses are submitted before GPU event reuse; next-generation source readiness proves prior host event waits have returned before re-recording.
- Partial abort records a fence for the submitted portion before workers are released; repeated finish does not add a fence. Existing issuer join and request stream synchronization remain.
- Defaults remain **1/1/1**. `STRATA_HIP_COPY_BATCH`, `STRATA_HIP_STAGE_BATCH`, `STRATA_HIP_STAGE_COPY_BATCH` are opt-in and guarded to HIP/Windows. Short prefill and decode retain their existing paths.
- The private production stager is shared with a small regression test. Existing `STRATA_PREFILL_TIMING` enables mechanical counts; no new production synchronization is added for diagnostics.

## Performance

Base: current `main`, **30ec18e**, engine 0.1.30. Same Release executable for both arms; **two independent processes per arm**, each with **2 warmups + 3 samples**. Main table is the median of six measured samples per arm; cold requests are excluded.

Qwen3.8-Flash-Next Q2_0; fixed 4,206 IDs / 4,205 processed by `Prefill::run`; cache 6,500; context 32,768; KV int8; MTP/spec4; auto request chunk 4,352; 1,575 borrowed slots; prompt reuse disabled. Windows / HIP 7.2 / RX 9070 XT / gfx1201.

| | Control 1/1/1 | Optimized 8/16/16 | Delta |
|---|---:|---:|---:|
| PP (DONE protocol) | 672.2 tok/s | 1,091.6 tok/s | +62.4% |
| `Prefill::run` wall | 6,144 ms | 3,738 ms | -39.2% |
| DONE prefill | 6,257.3 ms | 3,853.3 ms | -38.4% |
| TTFT | 6,281.5 ms | 3,890.5 ms | -38.1% |
| Copy-wait category | 2,746 ms | 536.5 ms | -80.5% |
| Minimum available RAM across process samples | 3.78 GiB | 4.28 GiB | |
| Maximum process working set | 32.38 GiB | 32.38 GiB | |

Independent repetition medians: **669.9 → 1,111.6 PP** and **672.8 → 1,087.7 PP**. RAM was sampled every 0.5 s with the unchanged 2 GiB availability watchdog. No TG improvement is claimed; these were prefill requests with one generated token.

Same payloads per request: **19,651 copies**, **13,938 direct pinned**, **5,713 staged**, **27,165,542,400 H2D bytes** (19,267,891,200 direct + 7,897,651,200 staged).

Warmed application-operation counters:

| | Control | Optimized |
|---|---:|---:|
| Direct `copied` records | 13,938 | 1,743 |
| Direct slot reuse waits | 13,938 | 1,743 |
| Staged `copied` records | 5,713 | 417 |
| Staged slot reuse waits | 5,713 | 417 |
| Host staging `dma_done` records | 5,713 | 358 |
| Host staging reuse waits | 5,697 | 5,697 |

These do not count all HIP API calls or measure hardware bus bandwidth. Compute-side consumption and `used` records are preserved.

## Correctness

Separate diagnostic A/B on this patch and current main, two identical requests per arm:

- **43,059,200 residual floats bitwise equal**, matching position records/SHA256, max absolute difference **0**, no NaN/Inf.
- **96 routing hashes equal** across the repeated requests; GDN fingerprints match and all **29,417,472 GDN state values/request are finite**. Fingerprints alone are not a full GDN state equality proof.
- Complete retained-cache readback before/after each request: **4,925 slots / 6,808,320,000 bytes**, exact.
- Complete borrowed/refill-cache readback: **1,575 slots / 2,177,280,000 bytes**, exact.
- Identical expert transfer counts/bytes and zero staging fence errors.
- `hip_prefill_stager`: **1,020 ownership cases** with byte-exact DMA readback, batches 1/2/4/8/16, job counts 1/15/16/17/31/32/33/65, rings 16/17, wraps, partial aborts, repeated finish and generations.
- Final clean build and selected CTest: hip_prefill_stager, hip_prefill_copy_groups, hip_intrinsics, hip_expert_cache_staging **4/4 pass**.

The additional, unmodified upstream `hip_handoff` test reports `separate copy/ring timeout` (exit 10) on this Windows/HIP system. Its target links only unchanged kernels/runtime, not prefill; this PR does not resolve that standalone doorbell issue. This is not a claim that the whole HIP suite passed.

Full residual/cache instrumentation and raw local results are excluded from the patch.

## Scope / limitations

Measured on **one RX 9070 XT**, **Windows HIP 7.2**. No CUDA/Linux/other AMD GPU performance or runtime-validation claim. Default-on elsewhere is not justified by this single-platform result.

No kernel, model, quality-setting, quantization, routing, cache-policy/size, MTP/spec, or decode changes. Current upstream GDN pipeline, MTP batching, gfx1201 dot4, session carve, doorbells and stream threshold are preserved. No research kernels, tracing tooling, model dumps or local paths are included.
## Copilot review follow-up

- 1339041 resets jobs, counters, submission state and generation even for empty staging lists. The production stager fixture now covers nonempty → empty → nonempty and passes 1,020 cases.
- 606f5d6 extracts the unchanged group planner and event contract into a private helper shared with production. The GPU fixture passes **216 generations** with direct, pinned-staged and forced-pageable sources, GPU rings 8/17/32, host rings 16/17, both batching settings enabled, repeated wraps and skipped/unrouted entries. An independent maximal-group oracle checks boundaries/window release; a held terminal used event verifies the actual reuse wait and terminal copied recording/publication. DMA payload readback is byte-exact. The issuer join guard now outlives its captured planner/bookkeeping on early return.
- End-to-end diagnostic A/B was repeated after this extraction: **43,059,200 residual floats bitwise equal**, max difference 0, finite; 96 routing hashes equal; GDN finite/fingerprints equal; complete retained and borrowed/refill cache checks exact in both arms; identical expert counts/bytes and zero staging fence errors.
- Additional matched performance check, same model/prompt/config and same executable, **one fresh process per arm, 2 warmups + 3 samples**: medians **723.0 → 1,162.2 PP tok/s**, Prefill::run **5,704 → 3,505 ms**, TTFT **5,844 → 3,641 ms**, copy-wait **2,508 → 496 ms**. This follow-up preserves the batching gain (+60.7%); it is separate from the original six-sample-per-arm table above and makes no new optimization or TG claim.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.