Issues / #1473
#1473 STRATA_SYCL_SPIN_MAX default: a JIT build of the B70 gets the A-series bound and loses 46-67% of decode
open · @demetree · 1 comments · View on GitHub
BenchmarksModels & quantsDocumentation
Description
### The 2,000,000 spin bound costs 46-67% of decode on a JIT build of the B70 0.1.40.3 made the device's spin bound on host flags a build option (`STRATA_SYCL_SPIN_MAX`) and defaulted it to 20,000 only for a `bmg` **AOT** build - 2,000,000 for everything else, on the grounds that the A-series JIT build needs it, where the GPU genuinely waits on the host's slow per-layer CPU expert work. A B70 built **without** AOT is not a `bmg` AOT build, so it inherits the A-series value. But the two cases are opposite: on the A-series the wait *ends* when the CPU expert work is done; on a card whose host handshake is not visible during a window, the wait is expected to run out and give up, and every extra read is time spent on nothing. `STRATA_VERIFY_NO_HOST=1` exists for exactly that case - the port's own note is that a kernel's writes to host-mapped memory are "not reliably visible while the graph runs". Same binary, one `-DSTRATA_SYCL_SPIN_MAX` apart, Arc Pro B70, OpenCL, Coder IQ1_M, `--spec 4 --mtp`, 128 greedy tokens, medians of two interleaved runs: | context | `-DSTRATA_SYCL_SPIN_MAX=2000000` | `=20000` | | |---|---|---|---| | 3-token prompt | 40.6 tok/s | **59.0** | +46% | | 2,048-token prompt | 25.6 tok/s | **42.9** | +67% | The loss grows with the context, which is what more expired waits looks like. With 20,000 the same engine decodes *faster* than 0.1.40.2 did (42.9 against 40.8 at the 2,048-token context; 59.0 against 51.0 short), so the whole of the regression I reported earlier is this default. ### What would help Keying the carve-out on the card (`bmg`) rather than on `STRATA_SYCL_AOT MATCHES "^bmg"` would cover the JIT build of the same card, which is what `docs/INTEL.md` tells people to avoid only for speed (the AOT build is the recommendation because of JIT time, not because the JIT build is wrong). Alternatively the bound could follow the property it exists for - whether this build's host handshake is visible during a window - which on the OpenCL backend it is not, on any card. No behaviour change either way in my measurements: `quantize_act_parity` and `sampler_parity` are 0 failures at both bounds, and through the API 2+2 answers 4.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.