Pull requests / #809
#809 Intel arc 0.1.39 + perf fixes
closed · @maxfridbe · 0 comentarios · En GitHub
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Descripción
Follow-up to #423 on top of the 0.1.39 release. Every change is in `sycl/` or the Intel docs; no shared upstream file is touched. ## 1. The 0.1.39 SYCL port did not build (fixes) Two upstream changes merged for the release did not reach the port's copies of the files they touch. On the B70 (oneAPI 2026.1, AOT `bmg-g31`) `sycl/` failed with 3 errors: - **#626 (thread affinity, `ec35cf0`):** `pin_current_thread` returns a `ThreadAffinity`, but `sycl/src/core/session.cpp` / `session.hpp` still held a `long long`. - **#559 (per-stage dense layers, `ba707d8` / `e50f266`):** `NativeDense::load` takes `layer_lo` / `layer_hi`, but the port's `native_dense.cpp` kept the old signature. With those two 3-way merged, 0.1.39 builds with 0 errors. On the card, 21 of 25 kernel tests pass, and benchy matches the 0.1.38 port's speeds. Of the four failures: - `iq_parity` and `ple_parity` need model fixtures. - `s2_expert_grouped_parity` covers the s2 path (it fails on the 0.1.38 port too). - `conversation_snapshot_test` aborted with `munmap_chunk(): invalid pointer`. The port's `conversation_state.cpp` (last synced at 0.1.31) and the test (0.1.32) predate #559's parking changes, while the port builds against upstream's `conversation_snapshot.hpp`, which #559 changed. Merged here as well. **Not ported yet:** #559's batch slots (`--batch` / `"parallel"`). The port's `verify.cpp` / `verify.hpp` do not have them. It is opt-in, so default use is not affected. ## 2. Prompt-path speedups Measured on a B70 with each model's served config. Decode speed is unchanged. - **Lend mirror.** The slots a prompt borrows have their experts re-read from the GGUF for every chunk and again for the refill. Those experts are now kept in pinned RAM and DMA'd. Same bytes, identical outputs. About 2 GB more RAM and about 2 s more at start. `STRATA_LEND_MIRROR=0` turns it off. - **QSA block scores as oneMKL GEMM tiles.** The old kernel used one sub-group per (query, block) pair, so its cost grew with the square of the context. Now the pooled keys are multiplied against the batch's indexer queries in fp32 tiles, followed by a small relu-sum kernel. 7-8x faster at 40K-256K, relative error 2-4e-7 vs FP64. `STRATA_SELECT_GEMM=0` restores the old kernel. - **Prompt attention scores one work-item per cell** (128-cell chunks), instead of a 5-step shuffle tree per head. 1.5x faster. Decode keeps the old kernel. `STRATA_ATTN_PERCELL=0` restores the old kernel. The last two are FP32 in a different summation order, so they are not bitwise, like the CUDA build's tensor-core paths. Greedy outputs follow the same text and part ways at a near-tie after tens of tokens. Time to first token, Coder IQ1_M on its served 256K config (measured on the 0.1.38 base): | Prompt | Before | After | |---|---:|---:| | 2,185 tokens | 4.9 s | 4.6 s | | 40K | 40.1 s | 33.6 s | | 128K | 140.9 s | 111.2 s | | 256K | 328.8 s | **235.1 s** | IQ2_XS and Swift 1.5 at 128K: 191 -> 156 s. Things tried that did not help are recorded in `sycl/TODO.md`: 15 attention variants, including two XMX versions, and expert grouping on the GPU (its "6.5%" was mostly the profiler's own cost). ## 3. benchy v1: a standard bench for Intel results `sycl/benchy.sh` runs every set-up model on fixed prompts (`sycl/bench/v1`: 20 / 2,185 / 8K / 40K / 128K / 256K tokens) with its own serve config, from a cold page cache. It reports: - speed: PP, TTFT, TG, drafts accepted; - resources: experts in VRAM and in the RAM mirror, VRAM, RAM, SSD reads, load time, power; - the system and each config's args. `docs/INTEL_PERFORMANCE.md` holds every measured number: current results, dated history, the B50 / B580 / 2x B70 results from #423, and "Submitting numbers", which asks testers to post benchy's report plus their config. ## 4. Smaller - **`setup_intel.py --check`** sized each model by the first card only. With two cards it now reports what fits across both. - **`strata-sycl.sh`:** with no `ONEAPI_DEVICE_SELECTOR` set, a `--layer-split` gets `level_zero:gpu` (the image pins `level_zero:0`). A set selector is still passed through. - **Duplicate `stage_room` fix removed:** where #423's fix and yours overlapped, this keeps yours. - **Docs:** an accuracy pass on `docs/INTEL.md` (build section, switches, a corrected "gather-bound" claim), plus the probe programs and three kernel benches (`attn_bench`, `sel_scores_bench`, `xmx_int8_bench`). ## Testing - The 0.1.39 build fix: built AOT for the B70, kernel tests run, benchy v1 on all three models. - This exact branch: a full B70 run (build, `ctest`, benchy v1) is in progress. I will post its report as a comment.
En el sitio
Enlaces a install, modelos, releases.