Pull requests / #1713

#1713 sycl/Arc: stepped verify window (opt-in STRATA_VERIFY_STEPPED) - CPU experts served between graph segments

open · @aslater3 · 0 Kommentare · Auf GitHub

BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsLinux

Beschreibung

Stacked on #1602 (Alchemist setup + staging-scratch fix). The first new commit here is the last one.

## What it does

`STRATA_VERIFY_STEPPED=1` is opt-in and off by default, so other GPUs are unaffected. It changes how the doorbell verify window serves CPU experts on Arc A-series.

Today the window is one graph. GPU kernels spin on host flags while the CPU pool serves the off-GPU experts inside it. On an A770 that spinning costs about 100–170 ms of GPU wait per window.

With `STRATA_VERIFY_STEPPED=1`:
- **Segmented capture.** The window is captured as graph segments, each ending at a layer's doorbell publish.
- **CPU experts between segments.** The host serves the CPU experts between segments, then raises the flag before it submits the segment that consumes them.
- **Pipelined helper.** Plan copy and VRAM groups run on a helper thread pinned to the host thread's SMT sibling. `STRATA_VERIFY_PIPE=0` and `STRATA_PIPE_HELPER_SIBLING=0` turn the pipelining and the pinning off.
- **Cached host buffers and plain payload stores.** In this mode no kernel reads a host store while it runs, and the host reads a segment's rings and payload only after that segment has drained. So the window's mapped buffers come from ordinary `malloc_host` instead of `host_malloc_polled`, and the doorbell publishes use plain payload stores instead of `sys_store_mapped`. The payload checksum still guards every host read.
  - `STRATA_VERIFY_HOST_UNCACHED=1` keeps both uncached.
  - With `STRATA_VERIFY_STEPPED` unset, every buffer and store is as on `main`.

## Measured

Setup:
- Hardware: Arc A770 16 GB (xe driver, oneAPI 2025.3, Linux 7.0), Ryzen 9 3900X, 32 GB DDR4.
- Model: Qwen3-Coder IQ1_M, learned24 expert profile, `--resident-experts --spec 4 --mtp`, greedy decoding, 256 new tokens.
- Prompt: code, 7,987 tokens.
- The runs are not interleaved. RAM headroom differs where noted, because resident experts are mlocked and the box has 32 GB.

| Build | Decode tok/s | GPU wait for the CPU per window |
|---|---|---|
| #1602 head (`712bbc9`) | 12.3 | 170 ms |
| this PR, `STRATA_VERIFY_STEPPED=0` | 12.6–13.1 | 134 ms |
| this PR, stepped, everything uncached (= `STRATA_VERIFY_HOST_UNCACHED=1`, measured on the dev branch) | 21.3–25.0 | 70–72 ms |
| this PR, stepped (dev branch, same commit) | **31.8–32.5** (4 runs) | 42.7–43.0 ms |
| this PR, stepped, 32K-token prompt (dev branch) | 27.2 | 41.9 ms |

The dev branch is this commit on top of an A770 rebase. Its resident set was 17.9 GiB, against 13.7 GiB for this branch with `STRATA_RESIDENT_HEADROOM_GIB=6`. On this PR branch alone, with the smaller resident set, stepped decode was 19.9 tok/s. It was memory-bound because part of the expert set was not resident.

Output check (greedy): this PR stepped, this PR stepped-off, and #1602 agree for the first 80 tokens. Run-to-run, #1602 itself agrees for 76–79 tokens on this prompt, so that is the existing nondeterminism, not a change.

## Not tested

- Arc A750/i915, B-series, and multi-GPU.
- The CUDA and HIP builds. The only shared-header change is one new SYCL-only declaration, `doorbell_plain_payload`, in `include/strata/kernels/elementwise.hpp`; its definition is in `sycl/`.

On exit, #1602 and this branch both sometimes segfault after the stats are printed (rc 139). It also happens without this commit.

Mehr auf der Site

Links zu Install, Modellen, Releases.