Pull requests / #1713
#1713 sycl/Arc: stepped verify window (opt-in STRATA_VERIFY_STEPPED) - CPU experts served between graph segments
open · @aslater3 · 0 comments · View on GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsLinux
Description
Stacked on #1602 (Alchemist setup + staging-scratch fix). The first new commit here is the last one. ## What it does `STRATA_VERIFY_STEPPED=1` is opt-in and off by default, so other GPUs are unaffected. It changes how the doorbell verify window serves CPU experts on Arc A-series. Today the window is one graph. GPU kernels spin on host flags while the CPU pool serves the off-GPU experts inside it. On an A770 that spinning costs about 100–170 ms of GPU wait per window. With `STRATA_VERIFY_STEPPED=1`: - **Segmented capture.** The window is captured as graph segments, each ending at a layer's doorbell publish. - **CPU experts between segments.** The host serves the CPU experts between segments, then raises the flag before it submits the segment that consumes them. - **Pipelined helper.** Plan copy and VRAM groups run on a helper thread pinned to the host thread's SMT sibling. `STRATA_VERIFY_PIPE=0` and `STRATA_PIPE_HELPER_SIBLING=0` turn the pipelining and the pinning off. - **Cached host buffers and plain payload stores.** In this mode no kernel reads a host store while it runs, and the host reads a segment's rings and payload only after that segment has drained. So the window's mapped buffers come from ordinary `malloc_host` instead of `host_malloc_polled`, and the doorbell publishes use plain payload stores instead of `sys_store_mapped`. The payload checksum still guards every host read. - `STRATA_VERIFY_HOST_UNCACHED=1` keeps both uncached. - With `STRATA_VERIFY_STEPPED` unset, every buffer and store is as on `main`. ## Measured Setup: - Hardware: Arc A770 16 GB (xe driver, oneAPI 2025.3, Linux 7.0), Ryzen 9 3900X, 32 GB DDR4. - Model: Qwen3-Coder IQ1_M, learned24 expert profile, `--resident-experts --spec 4 --mtp`, greedy decoding, 256 new tokens. - Prompt: code, 7,987 tokens. - The runs are not interleaved. RAM headroom differs where noted, because resident experts are mlocked and the box has 32 GB. | Build | Decode tok/s | GPU wait for the CPU per window | |---|---|---| | #1602 head (`712bbc9`) | 12.3 | 170 ms | | this PR, `STRATA_VERIFY_STEPPED=0` | 12.6–13.1 | 134 ms | | this PR, stepped, everything uncached (= `STRATA_VERIFY_HOST_UNCACHED=1`, measured on the dev branch) | 21.3–25.0 | 70–72 ms | | this PR, stepped (dev branch, same commit) | **31.8–32.5** (4 runs) | 42.7–43.0 ms | | this PR, stepped, 32K-token prompt (dev branch) | 27.2 | 41.9 ms | The dev branch is this commit on top of an A770 rebase. Its resident set was 17.9 GiB, against 13.7 GiB for this branch with `STRATA_RESIDENT_HEADROOM_GIB=6`. On this PR branch alone, with the smaller resident set, stepped decode was 19.9 tok/s. It was memory-bound because part of the expert set was not resident. Output check (greedy): this PR stepped, this PR stepped-off, and #1602 agree for the first 80 tokens. Run-to-run, #1602 itself agrees for 76–79 tokens on this prompt, so that is the existing nondeterminism, not a change. ## Not tested - Arc A750/i915, B-series, and multi-GPU. - The CUDA and HIP builds. The only shared-header change is one new SYCL-only declaration, `doorbell_plain_payload`, in `include/strata/kernels/elementwise.hpp`; its definition is in `sycl/`. On exit, #1602 and this branch both sometimes segfault after the stats are printed (rc 139). It also happens without this commit.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.