Pull requests / #1721
#1721 sycl: "parallel" works on the Intel server, and the request after a batch no longer answers token 0
open · @recutita · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIModels & quantsSecurityWindowsLinux
Beschreibung
## Title Issue: none. ## Summary `"parallel"` (the engine's `--batch` slots) did not work on the SYCL server: every request stopped after one token. 3d0ac1b fixes that in `sycl/serve/server_intel.py`. Once batch windows run, they exposed an engine bug on one card whose experts do not all fit VRAM: the next single request answered token 0 for every position (`!!!!`, about 1 tok/s). 30e892f fixes that and comes first, so no commit of this branch has `"parallel"` working but the request after it broken. ## What changed ### 30e892f: `sycl/src/program/generate.cpp`: with `STRATA_VERIFY_NO_HOST` and a mirror, adaptive swaps are off. - After the mirror is built, the existing `REFUSED` check (experts in neither VRAM nor the mirror under `STRATA_VERIFY_NO_HOST`) reads the switch once into `no_host`. A second check follows: if experts were missing from VRAM at start, `no_host` is on and the adaptive tier is on, `adapt_every` becomes 0 and one line says why. - Why: the mirror holds the experts missing from VRAM at the start. An adaptive swap moves an expert that was resident then out of VRAM, so it is in neither. The device-built plan sees it is not resident (`skip = 0`) and waits for the host's plan, which `NO_HOST` (the xe default in `strata-sycl.sh`) never sends; the wait gives up at its bound and the layer runs without a plan. With `STRATA_VERIFY_DEBUG=1`: 25 good windows, then `d_res` with experts that were resident at start set to -1, `skip[] = 0`, layer-0 router weights `-inf`, ids `0 1 2 ... 9`. - Why only after a batch window: the single-request path under `NO_HOST` never runs the host's per-layer service, so it never counts misses and the tier never swaps. A batch window's host path computes misses on the CPU and counts them. ### 3d0ac1b: `sycl/serve/server_intel.py`: `_pump` routes the batch slots' lines. - `install_logprobs()` replaces `StrataEngine._pump` (to keep `LP` lines beside their token), and its copy predates `--batch`: the slots' `BT` / `BDONE` lines went to the main queue, where no slot reads them. It now does what `serve/server.py`'s `_pump` does, in the same order: a line starting with one of `FATAL_PREFIXES` ends the engine; `BT` / `BDONE` lines go to their slot's queue; at EOF `last_err` keeps an `ERR` line and every slot queue gets `None`. The `LP` handling is unchanged. ## How it was checked Arc Pro B70 (32 GB, xe, Linux), AOT `bmg-g31`, Flash-Next IQ3_XXS on one card (12,545 expert slots in VRAM, the other experts mirrored), `--spec 4 --mtp`, 32K context, `"parallel": 2`. The same chat request (a BST class, 64 tokens, greedy) alone, twice at once, then alone again: | | alone | two at once | alone after the pair | |---|---|---|---| | main | not run | both stopped at 1-2 tokens (still waiting after 980 s; the engine log showed the batch done) | - | | 3d0ac1b only | right | right, 3.1-3.2 s | `!!!!...`, 48 s (76 s without `--mtp`) | | 3d0ac1b, `STRATA_VERIFY_NO_HOST=0`, no `--mtp` | right | right | right | | 3d0ac1b, `--adapt-every 0`, no `--mtp` | right | right | right, 2.1 s | | this branch | right, 2.0 s (same tokens as main) | right, 3.1-3.2 s | right, 1.4 s (same tokens as with `--adapt-every 0`) | - The single request's tokens are the same as main's: under `NO_HOST` the single-request path never swapped. - Each reply of the pair is one consistent text (no tokens of the other slot in it). The replies differ from each other and from the single request: on one card the batch windows compute part of the misses on the CPU, which rounds differently. With every expert in VRAM (two cards with `--expert-parallel-device`, a separate PR) all replies were the same. ## Extra Notes - The whole diff was reviewed by a human (the author). - Only the SYCL port is touched. - With every expert in VRAM 30e892f changes nothing (no misses at start); that is by the code, not run with this branch. With `NO_HOST` off the host serves such experts and swaps stay on. - The alternative to 30e892f, adding swapped-out experts to the mirror, would keep the tier on; it changes how the mirror is built and was not tried. - Tested on the B70 under xe on Linux only. Assisted-by: Claude Code with Opus 5.5 medium
Mehr auf der Site
Links zu Install, Modellen, Releases.