Pull requests / #868
#868 docs: Arc Pro B60 rows and notes, and the B60's PCI id (e211) in setup_intel.py
closed · @LocalXPU · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUModels & quantsDocumentation
描述
Results from two Arc Pro B60 24 GB cards, and the three things that were wrong or missing for that card in the Intel docs. ## What - `docs/INTEL_ARC.md`: two rows in "What has been run" (2x B60 with `--layer-split`, and one B60 with part of the experts mirrored in RAM), a short note on the B60 (PCI id, AOT target, `setvars.sh` under `set -u`, running without Docker, a misleading startup line on a layer split), and the host and software the rows were measured on. - `docs/COMMUNITY_BENCHMARKS.md` and `bench/results/2026-10-04-community-2x-arc-pro-b60/`: the report in the template's format, with the server configs, engine logs and per-request timings. - `sycl/setup_intel.py`: the B60 was listed as `e221`. Both cards here are `8086:e211` (`lspci -nn`), so setup called them `Intel GPU e211 (xe)` and guessed the VRAM from the BAR size. `e211` is added. `e221` stays: it is also a B60 (another B60 pair in this thread reports `8086:e221`), so setup now knows both ids. (An earlier version of this description guessed from the kernel's id list that `e221` was not a B60; that was wrong.) `tools/test_setup_sycl.py` has two new checks (8 tests pass; the new one fails without the table entry). ## Numbers 2x Arc Pro B60 24 GB (456 GB/s each, PCIe 3.0 x8), Ryzen 5 5600, 64 GB RAM (62 GiB usable), Ubuntu 24.04, kernel 6.17 (`xe`), NEO 26.09.37435.12, oneAPI 2026.1.1. `6f32ec0` plus the compile fix and the ring-wait fix (separate PRs), AOT `bmg-g21`. Context 8,192, KV int8, `--spec 4`, greedy, one request at a time; every expert resident on both cards: | | decode | prompt (2,129 tokens) | | --- | --- | --- | | Coder IQ1_M, `--layer-split 24` | 56.6 tok/s | 428 tok/s | | Flash-Next IQ2_XS, `--layer-split 25` | 58.6 to 61.3 tok/s | 436 tok/s | Decode and prompt speed are the server's own `timings` per request. After 1.9K tokens of code context decode was 57.2 (Coder) and 55.5 (IQ2_XS). One B60, Coder IQ1_M, 4,042 of 12,288 experts mirrored: 11.9 tok/s. The report has every request, draft acceptance, VRAM per card, and what was not tested (streaming, tools, contexts above 2.2K). ## Notes - The `INTEL_ARC.md` sentence that nobody has run the 0.1.39 port on an Arc is now out of date for the B70 (#809) and for this report; I left the maintainers' sentence alone and said what ran where. - `docs/COMMUNITY_BENCHMARKS.md` asks for results-only submissions to stay separate from engine changes. This one has three lines in `setup_intel.py` and a test; I can split them out if you prefer. - #797 rewrites `docs/INTEL_ARC.md` and is marked do-not-merge; #809 touches `docs/INTEL.md` and `setup_intel.py` in a different place. I expect this to merge cleanly with #809 and to need a rebase over #797 if that one lands.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。