Pull requests / #1710

#1710 sycl: experimental A770 fast router and paired mapped transfers

open · draft · @duncanmcqueen · 0 comentarios · En GitHub

BenchmarksSetup & installAMD / HIPModels & quants

Descripción

## Summary

Experimental follow-up to #1602: a fast SYCL router, paired uncached mapped-memory transfers, and an opt-in host-clock GPU-stage profiler for investigating slow Arc A770 decode.

**Dependency:** #1602 is still open. This branch includes its eight commits plus the speed-work commit `501f675`, so it builds against upstream `main`. The new work is the seven-file diff from `duncanmcqueen:intel-sycl-handshake` to `duncanmcqueen:a770-speed-upstream`. Once #1602 merges, this follow-up should be rebased to remove the dependency commits from its review diff.

**Draft status:** normal inference and the targeted parity tests pass, but the experimental profiler still has unresolved synchronization concerns and enabled runs have encountered a payload-arrival timeout. This is submitted for review and further investigation, not as a claim of a production-ready profiler or 20 tok/s decode.

## What changed

- Port the fork's fast router into the base SYCL engine. A tree sum brackets the original serial reciprocal, with an ascending serial fallback when float rounding is ambiguous. Selection and renormalization run in subgroup 0. `STRATA_SYCL_ROUTER_OLD=1` keeps the old implementation.
- Add a router comparison probe covering 82,944 rows across four geometries, including ties, NaN tails and the forced serial fallback. Compare expert IDs and weights bitwise against the existing router.
- Publish activation payloads as paired 64-bit uncached stores, preserving checksum word positions and scalar fallback for unaligned pointers and odd tails.
- Read mapped CPU results using paired 64-bit uncached transactions. Add a probe for repeated rings, all three publish variants, odd sizes, unaligned fallback and mapped readback.
- Register both probes with CTest.
- Replace the inert zero-timestamp stage profiler with the fork's host-clock approach: an opt-in host writer thread updates a host USM word read uncached by timestamp kernels. The writer is stopped and joined at teardown. This diagnostic remains experimental.

## Measurements

Hardware: Intel Arc A770 16 GB (`xe`), Ryzen 5 5600, 31 GiB RAM, oneAPI 2026.1. The A770 is behind the chipset's PCIe 3.0 x4 connection; Strata's host-to-device probe measured 1.7-1.8 GB/s.

Model: Qwen3.8-Flash-Next Coder IQ1_M. Engine based on upstream 0.1.41. Tests retained 131,072-token context capacity, INT8 KV with 32,768 resident cells, MTP speculation, and disabled prompt caching. Actual prompts were 73-546 tokens; these are not full-128K-context measurements. Generated-token rates include reasoning tokens.

- For the Coder's 256-expert/top-10 router shape, host submission plus execution fell from **395.47 to 118.62 us per single-token call**, with bitwise-matching outputs.
- The earlier pinned-mirror configuration decoded at about **2.5 tok/s**. Most of the end-to-end improvement came from **configuration changes**: CPU-computed misses, a larger VRAM cache, and smaller prompt buffers. It should not be attributed entirely to this code patch.
- With CPU-computed misses, cache borrowing, an **8.98 GiB VRAM expert cache**, and the paired-transfer patch, three 546-token-prompt runs generated 128 tokens at **13.5, 11.9 and 12.3 tok/s**. The corresponding fresh-prompt rates were **56.9, 57.0 and 57.5 tok/s**. Zero prompt tokens were reused.
- A warmed 73-token prompt generated 128 tokens at **11.4 tok/s**.
- The original 542-token arithmetic regression check returned final content **`41`**. Cold first-request generation remains much slower because of initialization/JIT costs.
- The **20 tok/s goal has not been achieved**.

Useful tested configuration changes relative to the pinned-mirror setup:

```text
remove: --stream-experts
add:    --resident-experts --pool-workers 7 --pool-affinity auto
        --adapt-swaps 8 --prefill-borrow --no-spec-split
set:    --prefill 1024 --vram-reserve-mib 300
keep:   --max-context 131072 --kv int8 --kv-resident 32768
        --spec 4 --spec-min-p 0.5 --mtp <existing MTP path>
env:    STRATA_VERIFY_NO_HOST=0
```

The existing wrapper consumes `STRATA_VERIFY_NO_HOST=0` to disable its default NO_HOST setting. On `xe` without `--stream-experts`, the engine defaults the PCIe miss share to zero, so CPU workers compute misses from the resident RAM complement. Adequate free RAM is still required.

## Validation

Rechecked against the latest fetched upstream `main` (`fb58e0d`); the branch was already based on it and needed no conflict resolution.

- Build: `strata`, `sycl_router_fast`, `sycl_payload_pairs`, `elementwise_parity`, `router_top10_parity` succeeded with oneAPI.
- A770 CTest: **4/4 passed**, 7.70 seconds:
  - `sycl_router_fast`
  - `sycl_payload_pairs`
  - `elementwise_parity`
  - `router_top10_parity`
- `git diff --check` passed.
- Quality review found no blocking defect in the unprofiled router or paired-payload changes.

## Known profiler limitations

- Runs with `STRATA_VERIFY_PROFILE=1` encountered `verify: layer ... rang but its payload never arrived whole`. The cause is unresolved; ordinary unprofiled benchmarks passed.
- Review flagged the clock word's concurrent atomic host stores and ordinary non-atomic uncached device reads as lacking a correctness-guaranteed atomic synchronization contract. Cache-bypass hints alone do not establish atomicity. This has **not** been established as the cause of the payload timeout.
- The host clock writer consumes CPU resources and timestamp kernels add overhead. Profile-mode throughput is diagnostic and should not be used as the normal-inference speed measurement.

En el sitio

Enlaces a install, modelos, releases.