Issues / #1054
#1054 [SYCL / Intel Arc] 2x Arc Pro B70 --layer-split auto deadlocks at startup after the host-mirror fill
open · @djbrettb · 3 comentários · No GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation
Descrição
**Summary.** On two Arc Pro B70s, Flash-Next IQ3_XXS and IQ3_S with `--layer-split auto` load, fill the first card's expert cache, mirror the remaining experts into pinned host memory, and then never reach "session is up". One host thread spins at 100 % CPU, the second card reports 100 % busy at ~88 W (a spin, not compute: real compute on this card draws 200+ W), and the first card idles at 0 MHz. The server keeps printing `still starting (...)`. It happens with and without PR #866 and with `STRATA_VERIFY_NO_HOST=0` or `=1`. The same build runs IQ2_XS on **one** B70 without any problem: 79 tok/s short decode, a 12-minute soak with 156/156 requests and 0 xe engine resets. ### What looks inconsistent in the log The split planner says both caches together hold every profiled pair. But the session's expert cache is sized only on CUDA0, and the rest goes to the host mirror, while CUDA1 still has ~28 GiB free: ``` strata generate: layer split across 2 GPUs: CUDA0, then CUDA1 (split auto) strata generate: PCIe probe: 18.2 GB/s host->device (best of 18.0 18.2 18.1 18.1) -> pcie_frac 0.50 (default 0.55) strata generate: layer split: CUDA1 PCIe probe 21.1 GB/s (best of 21.1 19.2 18.0 17.9) -> pcie_frac 0.55 strata generate: layer split: CUDA1 holds its weights; 28.57 GiB free (its session follows the split search) strata generate: layer split auto: CUDA0 32 SMs at 2.80 GHz -> 0.81 ms per layer, 24.87 GiB free before its session carve strata generate: layer split auto: CUDA1 32 SMs at 2.80 GHz -> 0.81 ms per layer, 23.81 GiB free before its session carve strata generate: layer split auto: K=29 - predicted 38.9 ms per decode window; the caches hold 24576 of 24576 profiled pairs (~100.0% of the routed mass) strata generate: layer split: CUDA1 holds its weights, session [29, 48) and the head; 27.92 GiB free strata mtp: draft layer loaded, 839 MiB of VRAM (experts 675, dense 111), files read in 0.39 s (2005 MiB/s) strata generate: expert cache auto: 28.22 GiB free, 2048 MiB reserved (+0 MiB for the draft head) -> 11401 slots strata generate: expert cache 14848 slots, 21.74 GiB of VRAM; policy is strata generate: pre-filled 14848 of 14848 slots from the profile; slot 0 verified strata generate: 9728 of 9728 experts missing from VRAM mirrored in pinned host memory (18.23 GiB, 2.7 s); the GPU reads them over PCIe (nothing further; the server prints "still starting (N s)" until killed) ``` IQ3_S behaves the same way: K=26, 13,312 slots / 24.19 GiB on CUDA0, and 11,264 experts (22.64 GiB) mirrored. `--check` (`sycl/setup_intel.py --check`) reports IQ3_XXS and IQ3_S as "fits in VRAM across the 2 cards (layer split)". So we expected no host mirror at all. ### State while stuck - Engine process: 78 threads. 77 sleep in `futex_do_wait`; one runs at 100 % CPU. - Card 0 (`0000:03:00.0`): 0 % busy, 0 MHz. Card 1 (`0000:08:00.0`): 100 % busy (gtidle residency), 2.8 GHz, ~88 W. - No xe engine resets or AER errors while it hangs. - Killing the container produced one reset on card 1: `xe 0000:08:00.0: [drm] Tile0: GT0: Engine reset: engine_class=ccs, logical_mask: 0x1, guc_id=2, state=0x289` followed by `Fault response: Unsuccessful -ENOENT`. The card was healthy afterwards; llama.cpp loaded and served on it. Our guess, which we have not verified in the source: with the layer split, the second stage waits on a host-mapped flag/doorbell that the host writes while CUDA1's graph is running. `docs/INTEL.md` notes that host-to-device visibility during a kernel is unreliable on this platform. The host then waits for CUDA1, so the two never meet. ### What we tried | variant | result | |---|---| | `maxfridbe/Strata_B70@b70` 8dd4aaa (AOT `bmg-g31`), IQ3_S and IQ3_XXS, `--layer-split auto`, `--vram-reserve-mib 2048` | hang (both) | | same + PR #866's change applied by its own `fixups.py` regex (2 sites) | hang | | same + PR #866 + `STRATA_VERIFY_NO_HOST=0` (made overridable in `strata-sycl.sh`) | hang | | IQ2_XS on ONE card (no split), `--vram-reserve-mib 3072` | works: 79 tok/s short, 66 tok/s at 2.5K prompt, 12-min soak clean | | llama.cpp b11151 SYCL, same IQ3_XXS GGUF, `-sm layer` across both cards, `--override-tensor per_layer_token_embd=CPU` | works: ~25 tok/s, 10-min soak clean | Not tried yet: `--layer-split K` with an explicit K, `--split-device`, `STRATA_VERIFY_DEVICE_PLAN=0`, `STRATA_WARM_GRAPHS=0`. We can run any of these, or a debug build, on request. ### Environment - 2× Intel Arc Pro B70 32 GB (ASRock boards, subsystem 1849:6025), each in its own CPU root port, **PCIe 5.0 x8 each** (the x16 slot is bifurcated x8x8 through a C-Payne MCIO Gen5 retimer). Both cards are power-capped at 230 W (`power1_cap`). Separate IOMMU groups, 32 GB ReBAR each. - AMD EPYC 4464P, 128 GB DDR5 ECC; ASRock Rack B650D4U-2L2T/BCM. Proxmox VE 9.2 (Debian 13), kernel 7.0.14-20-pve, xe driver, GuC 70.72.1 / HuC 8.2.10 / DMC 2.6. - Engine image: Ubuntu 24.04 with Intel oneAPI 2026.1.1 (`icpx 2026.1.1.20260724`) and the compute runtime from the kobuk-team PPA: `intel-opencl-icd` / `libze-intel-gpu1` / `intel-ocloc` 26.31.39395.14. Built with `-DSTRATA_SYCL_AOT=bmg-g31 -DSTRATA_SYCL_PARITY=OFF`. - Engine args (paths shortened): `--pack packs/iq3_xxs --native …IQ3_XXS-00001-of-00002.gguf --ple-gguf …IQ3_XXS-00002-of-00002.gguf --expert-profile data/expert-profile.bin --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp mtp/rt --max-context 32768 --kv int8 --stream-experts --vram-reserve-mib 2048 --layer-split auto` `ONEAPI_DEVICE_SELECTOR=level_zero:gpu` (from `strata-sycl.sh`), `STRATA_VERIFY_DEVICE_PLAN=1`. ### Related - #867 (2x B60, host mirror, `never rang (graph finished)`) was fixed by #866. Here #866 is applied and the startup still hangs: there is no error, just a spin. So it looks like a different path. - The v0.1.39 SYCL build failure we hit is already #891, so we built from the port author's `b70` branch. cc @maxfridbe. Thanks for the port: the single-card numbers on the B70 are excellent, and we'd love to run IQ3 across both cards.
No site
Links install, modelos, releases.