Pull requests / #1111
#1111 Intel arc 0.1.40
open · @maxfridbe · 0 Kommentare · Auf GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
Replaces #809, which closed when `main`'s history was rewritten. Rebuilt on the new `main`: the 18 commits
of #809 cherry-picked, plus one commit that refreshes the port to 0.1.40. Nothing outside `sycl/` and
`docs/INTEL*.md` changes.
## Why the refresh was needed
`main`'s SYCL build had stopped compiling: `fused_gr`, `expert_source` and `prefill/gemm` no longer
matched the 0.1.40 headers, because `sycl/` still held its 0.1.39 copies.
The port is re-migrated with dpct and 3-way merged against 5047172, the commit it was migrated from.
That leaves 175 conflicts and three hand-ported files (`verify.cpp`, `mtp.cpp`,
`native_expert_parity.cpp`), now resolved. dpct's mistranslations in the new code are listed in
`docs/INTEL.md` (0.1.40).
## Fixed with it
- **#866:** the per-layer ring waits read `DPCT_CHECK_ERROR(q->ext_oneapi_empty())`, which is always 0.
So whenever experts were mirrored and `STRATA_VERIFY_NO_HOST` was off, any layer slower than 2 ms
was reported as "never rang".
- **#1054:** with a layer split, the host mirror now holds only the first GPU's layers.
- Before, it was built before the later stages' caches existed, in the first GPU's context. Every
later-stage expert was mirrored (13-18 GiB of RAM), then the later GPU's cache fill copied from
that pinned memory. Using it across contexts is undefined in SYCL; on two B70s it hung at startup.
- This doesn't reproduce on a B70 + B65, so the fix is untested on the reporter's hardware.
## Port-side choices
- **Shared-expert side stream:** `STRATA_SH_STREAM` is off. A recorded SYCL graph only takes its
own queue's work.
- **MTP drafter:** keeps launch-then-wait instead of host spins on mapped outputs.
- **Expert kernels:** upstream's warp-shaped kernels are opt-in (`STRATA_EXPERT_SPLIT=1`); the port's
grids stay the default.
- **VMM:** `vmm.cpp` is upstream's no-VMM branch, so `--kv-grow` and `--vram-elastic` stay off.
- **Not built:** `gemm_bf16_parity` (cuBLAS) and `native_expert_bench` (raw CUDA).
- **Version:** the engine reports 0.1.40-sycl, which `setup.py`'s `MIN_ENGINE` requires.
## Tested (B70, plus a B70 + B65 split)
- **ctest:** 34 of 34 pass, with no engine resets. `mmvq_multi_parity`'s negative control has no power
on the port's wide kernels, which use one layout at every column count; the test reports that the way
it reports gfx906's wave64 layout.
- **Output:** IQ2_XS greedy output is identical to the 0.1.39 port. A `--layer-split 29` run over
B70 + B65 matches one card token for token.
- **Decode** (short chat prompt, 64 tokens): 63.5 -> 68.0 tok/s on one card, 57.9 -> 62.7 tok/s on
the split.
- **Conversation cache:** with streamed K/V, two conversations sharing a 15K-token prefix took turns
(park, mount the shared checkpoint, restore). Every reply was sane, with `STRATA_VERIFY_NO_HOST=1`
and with per-layer waits. The B60 token-0 report from #809 did not reproduce.
Mehr auf der Site
Links zu Install, Modellen, Releases.