Pull requests / #797

#797 sycl: enable Arc A770 inference and document measured performance

closed · @atalantisalt03 · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDASecurityDocumentationLinux

Descripción

!!!DO NOT MERGE!!! (This describes my experience on an a770 and getting it to work, not polished code. GTP-6-Astra did everything since I don't know sycl at all lol. Use whatever is good from this if yall want, or just close it)


All code changes by gpt 6 astra. When I learned about strata I tasked astra with getting it working, it wasnt able to (see the 3tps result), after https://github.com/Niko1221/Strata/releases/tag/v0.1.39 I had it retry, and it sort of works. (Not verified for production at all).

## Problem and result

The v0.1.39 SYCL engine could not reach native decode on an Arc A770 16 GB: migrated shared-API signatures did not compile, unavailable free-memory telemetry prevented a usable cache, large cache/mirror allocations failed, and the native router hung at a divergent work-group barrier.

This branch gets the existing IQ2_XS model running on that card. With requested cache3000, median decode is **17.8 tok/s after 511 prompt tokens and 16.6 tok/s after 4096 prompt tokens**. This is still experimental and approximately 21% below the preserved RTX 5060 Ti CUDA cache3000 result. **Draft for author review; not ready to merge.**

## Changes

- Refresh the migrated `ThreadAffinity` and `NativeDense::load` interfaces.
- Fall back to UUID-matched Level Zero Sysman telemetry when the SYCL free-memory extension is unavailable; never substitute total VRAM for free memory.
- Use relaxed Level Zero allocations and matching frees for oversized expert-cache and pinned-mirror buffers.
- Replace the native router's work-group barrier with a subgroup barrier: the other seven subgroup rows have already returned. Add a 20-case regression/parity test and a migration fixup.
- Reject no-host verification without a device plan or with incomplete streaming mirror coverage. Retain opt-in eager diagnostics used to locate the hang; diagnostics are disabled in the measured runs.
- Update `docs/INTEL_ARC.md` and link it from `docs/INTEL.md`, including required A770 flags, memory requirements, measurements and limitations. Add a portable benchmark JSON with the saved corpus, per-request rates, output hashes and provenance.

## Measurements

A770 16 GB, Ryzen 5 3600, Linux VM with approximately 94 GiB DDR4 RAM; NEO 26.35.39758.10, oneAPI 2026.1.1 and oneMKL 2026.1. SPIR-V/JIT. Cache3000 becomes 3134 slots / 4.22 GiB; all 21,442 non-VRAM experts are mirrored in 28.80 GiB of pinned host memory.

| Configuration | Short/long prefill tok/s | Short/long decode tok/s |
|---|---:|---:|
| A770, FP64 emulation enabled | 116.5 / 337.7 | 17.8 / 16.6 |
| A770, FP64 emulation disabled | 116.5 / 337.6 | 17.8 / 16.6 |
| Preserved CUDA cache3000 | 83.8 / 345.6 | 22.5 / 21.1 |

One warmup plus three trials per prompt size, 256 generated tokens per request, zero prompt reuse. Both FP64 arms return exactly matching text for all seven requests. They ran sequentially, not randomized. The older custom A770 port measured 2.7/2.7 tok/s on the same corpus/cache setting; the new release, kernels, host mirror and CPU policy all changed, so this is not a single-fix speedup attribution or an XMX claim.

## Validation and review limits

The engine changes were compiled and tested on the A770 before the documentation commit:

- 5/5 targeted GPU CTests pass, including 20 native-router cases with exact IDs and independent CPU softmax references.
- Real-model expert checks at layers 0/1/3/47 pass their 3% reference tolerance.
- A 5 GiB device + 5 GiB host allocation probe verifies kernel access beyond 4 GiB and memory release.
- Two fresh-process math controls per FP64 setting return identical correct token IDs; both seven-request throughput runs complete and clean up.
- **Strict API correctness remains 4/6.** `sum(i*i for i in range(4))` returns 30 instead of 14 twice; the same failure exists in prior CUDA/custom A770 evidence. No full-model quality or numerical-parity claim.
- The earlier broad suite had unresolved fixture, timeout, bitwise-comparison and snapshot-segfault failures; it was not rerun after these fixes. B70/other Intel hardware was not retested.
- Corpus/request hashes, response hashes, documented medians and source provenance were checked; `git diff --check` passes.

The general A770 configuration keeps FP64 emulation enabled and requires `SYCL_PROGRAM_APPEND_COMPILE_OPTIONS=-ze-intel-greater-than-4GB-buffer-required`. Disabling emulation is qualified only for the tested native greedy path.

En el sitio

Enlaces a install, modelos, releases.