Issues / #1662
#1662 Intel Arc B60 SYCL: v0.1.39 FastFix reaches 25.3 tok/s vs 10-12 tok/s in our v0.1.40.3 integration
open · @Goldlionren · 0 comentários · No GitHub
BenchmarksSetup & installMulti-GPUAMD / HIPModels & quants
Descrição
## Summary We investigated a large **effective Decode throughput difference on one Intel Arc Pro B60 (24 GiB)** while integrating Strata's SYCL backend with a custom, chunked pinned-host GGUF expert mirror for Swift-Qwen3.8-Flash-Next IQ3_XXS. - Our **v0.1.40.3-based JR integration** was functionally stable, but typical interactive/sustained Decode was **~8–12 tok/s**. - A **v0.1.39-based JR FastFix**, retaining the historical high-residency/adaptive execution path and incorporating targeted correctness and memory-lifetime repairs, completed **2,048 output tokens at 25.3 tok/s**, Adaptive ON, MTP Spec 4, with no observed readiness timeout, `!!!!!` corruption, or new GPU fault/reset in that bounded test. **Scope limitation:** This is **not** a controlled apples-to-apples benchmark of unmodified upstream v0.1.39 versus v0.1.40.3. Our JR versions differ in cache budget, runtime (Docker/native), and custom integration code; an exact matched historical baseline was blocked by an older binary's GPU startup fault. We are reporting measured results and actionable observations, **not claiming a particular upstream commit caused a 2–3× regression**. ## Hardware and versions - Intel Arc Pro B60 24 GiB, Battlemage BMG-G21, PCI ID `8086:e211`, PCIe Gen4 ×4; single GPU used for inference. - Ubuntu 24.04, `xe` driver, kernel `7.0.0-34-generic`; approximately 62 GiB usable host RAM. - Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS, same frozen two-part GGUF model, native pack, MTP and original ranking. - Both final implementations use the **oneAPI 2026.1 family**. The faster FastFix runs in a preserved Docker user-space image (runtime 2026.1.0, UR 0.12.0, Level Zero loader 1.32.0). - The original unpatched historical binary produced long-output `root!!!!!` corruption. **The successful FastFix is a repaired binary**, not that original binary. ## Measurements | | JR v0.1.40.3 RC1 | JR v0.1.39-based FastFix | |---|---:|---:| | GPU resident experts | 9,248 / 14.99 GiB | 10,831 / 17.57 GiB | | Other experts pinned-host mirrored | 15,328 / 15,328 | 13,745 / 13,745 | | VRAM reserve | 1,536 MiB | 768 MiB | | Adaptive activity | No net residency changes observed in a diagnostic sample | `--adapt-every 4 --adapt-swaps 96`, active | | MTP | Spec 4 | Spec 4 | | Context | 32K default; 128K tested separately | 128K configured; actual 120K input not re-qualified | | Decode observation | ~8–12 tok/s; 512-token response 10.86 tok/s | **25.3 tok/s over 2,048 emitted tokens** | The FastFix 2,048-token output hit the requested output cap, so it establishes bounded inference stability, not the completeness of a long-form document. Separate FastFix English/Chinese/math follow-ups returned 21.2/17.2/11.2 tok/s. This is not a guarantee of 25 tok/s for every prompt or proof of long-term endurance. ## Findings relevant to SYCL maintainers ### 1. Device-only routes can leave adaptive heat feedback incomplete With `STRATA_VERIFY_NO_HOST=1`, the device-planned verifier bypasses ordinary host expert dispatch in our v0.1.40.3-based branch. Our device route capture showed **53.648% of verifier route entries using host mirror addresses**, including speculative input rows. This is *not* measured PCIe traffic. Host-dispatch expert-heat counters may not observe all of these routes, which could reduce adaptive residency's ability to learn real Decode access patterns. A focused graph-visible native expert microprobe (actual IQ3_XXS weights, two tested layers, grouped T=1/4/6) measured **13.9–32.4× longer operations on pinned-host than device-resident weights**. That is a microbenchmark ratio, not whole-engine speedup. At the *same* engine and expert-cache byte budget, an offline-selected alternate ranking improved several short fixtures substantially, but the independent long-form fixture regressed ~16.5%: residency is a major lever and is workload-sensitive. ### 2. Custom pinned GGUF mirror initially did not support dynamic eviction Our **JR-specific** chunked pinned-host mirror populated entries only for experts *initially* missing from VRAM. Adaptive swaps could later evict an originally GPU-resident expert (`residency=-1`) without creating/updating its GGUF mirror pointer. Device plan coverage could then fail and revert to a slower host-assisted path. This flaw is **confirmed in our custom integration, not proof of a defect in unmodified upstream**. We also discovered a per-group x/IDs/weights payload-lifetime hazard when GPU progress ran ahead of the host loop. We fixed mirror slot ownership/publication and per-layer payload lifetime in separate changes. During a successful final Adaptive ON test there were **14,066 swaps / 251 generations and zero coverage holes**, no ring timeout, no new fault/reset, and 2,048 emitted tokens at 25.3 tok/s. We did *not* isolate the exact individual contribution of each patch. Earlier failed traces had A-plan readiness timeouts (ring 33 or 47, e.g. ring47/layer46/A observed40), but lacked the precise failed expert ID. We cannot prove those specific historical failures were exclusively caused by the mirror gap. ### 3. oneAPI 2026 and PLE I/O are not proven as the main slowdown The faster, repaired Docker build also uses oneAPI 2026.1; hence an Intel oneAPI 2026 upgrade alone cannot explain our observed difference. Separately, correcting PLE staging and changing the PLE table from Direct I/O to mmap produced a modest ~10.8→11.5 tok/s improvement in one comparison, but did not account for the whole performance difference. ## Questions / possible upstream improvements 1. In the SYCL **device-only verifier** path, is there a supported route/heat feedback mechanism for the adaptive expert scheduler? Are there specific patches newer than v0.1.40.3 worth evaluating? 2. For partially resident **streamed GGUF experts** with Adaptive ON, what invariant maintains valid host backing, completed copy, and generation-safe residency/mirror metadata across captured verifier graphs? Does the current upstream GGUF implementation guarantee coverage for newly evicted experts? 3. Is there recommended low-overhead instrumentation for **warm SYCL captured-graph** execution and pinned-host expert work on B60? Our PTI collection failed to expose warm native kernels. 4. What is the recommended B60 configuration/implementation to preserve high Decode throughput *and* the v0.1.40.x correctness/synchronization fixes? **Engineering choice:** we froze the successfully repaired v0.1.39-based FastFix for our B60 deployment rather than adopt our slower v0.1.40.3 integration. This is a local, hardware/workload-specific decision and **not a conclusion that all upstream v0.1.40.x builds are slower**. ## Public references - [JR-Strata FastFix frozen production release](https://github.com/Goldlionren/JR-Strata/releases/tag/jr-b60-sycl-fastfix-prod-20261009), including exact engine ELF, source/tag and checksums. - [Production identity / known qualification limits](https://github.com/Goldlionren/JR-Strata/blob/main/JR_STRATA_SYCL_PRODUCTION_FREEZE.md). - [Clean-install and external model/runtime requirements](https://github.com/Goldlionren/JR-Strata/blob/main/JR_STRATA_SYCL_CLEAN_INSTALL_SOP.md). - Related discussion: #867 (B60 host-mirror/ring waits), #1397 (SYCL version-specific performance on B70, different backend), #1302 (B65 SYCL expert-cache fill and reset). Thank you for maintaining Strata and the Intel SYCL port. We can provide a narrowly scoped source diff and sanitized diagnostic logs if useful.
No site
Links install, modelos, releases.