Pull requests / #409
#409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory
closed · @eelgaev · 0 comentarios · En GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Descripción
## aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory This PR makes Strata build and run on aarch64, and `./setup.sh` work end to end on an NVIDIA DGX Spark. It is experimental and community-tested, in the same spirit as the AMD HIP and sm_70 builds: **one DGX Spark, Qwen3.8-Flash-Next IQ2_XS installed by setup**. x86 builds are unchanged. ### Speed (IQ2_XS, setup's own `run-iq2_xs.sh`: 128K context, `--kv int8`) All 24,576 experts fit in the GPU's expert cache (33.0 GiB), so the CPU computes none. | Prompt | Context | Prefill | Decode | MTP acceptance | | ---: | ---: | ---: | ---: | ---: | | 2K | 0 | 928 tok/s | 58.7 tok/s | 78% | | 8K | 0 | 1,356 tok/s | 62.0 tok/s | 81% | | 2K | 16K | 1,421 tok/s | 60.5 tok/s | 83% | | 8K | 16K | 1,450 tok/s | 58.2 tok/s | 79% | | 2K | 64K | 1,505 tok/s | 58.9 tok/s | 80% | | 8K | 64K | 1,515 tok/s | 54.9 tok/s | 75% | llama-benchy 0.4.0, 3 runs after a warm-up, `--no-cache`, one request at a time, 128 tokens decoded per request. Prefill and decode are from the server log (`strata serve: prompt ...`). benchy's `pp` column cannot time Strata's server, because the first stream chunk is sent before the prompt is read; details are in `docs/DGX_SPARK.md`. ### What changes The change splits naturally into three parts. Happy to send them as separate PRs if you prefer. **1. Build and engine on aarch64** - `src/kernels/cpu/portable.cpp` replaces the AVX-only kernel files on non-x86 (CMake chooses by `CMAKE_SYSTEM_PROCESSOR`). It contains the activation quantizer (expert.cpp's scalar loop) and the BF16 router dot (NEON). The AVX multi-token kernels report themselves unsupported, so native experts use ggml-cpu's NEON `vec_dot`. - Q2_0 expert rows (the down projections of every setup model) go through ggml-cpu on non-x86 (`q2_own_kernels()`). x86 keeps `q2_rows_any`. - `include/strata/platform/cpu_relax.hpp`: `_mm_pause` / `_mm_sfence` on x86 (as before), `yield` / `dmb oshst` on aarch64. - `-ffp-contract=off` for C/C++ on non-x86 GCC/Clang, as an x86 build without `-mfma` effectively has, so host roundings match x86 and ggml. (`quantize_act_parity`'s Q8_K case failed without it.) - `GGML_CPU_ARM_ARCH` turns `GGML_NATIVE` off: gcc's `-mcpu=native` finds no features on GB10's Cortex-X925/A725, which leaves ggml-cpu without dotprod/i8mm. **2. Expert cache sizing on unified memory** - `device_free_bytes()` is used by both the cache sizing and `ExpertCache::open`'s check. When the driver reports that the GPU shares host memory (`cudaDevAttrIntegrated`), it counts `MemAvailable` less 6 GiB (`STRATA_UMA_HEADROOM_GIB`) as free. `cudaMemGetInfo` counts only free RAM, and with `--mmap-experts` the page cache is mostly the GGUF's own pages. NVIDIA's DGX Spark Porting Guide (§5.5) recommends the same; swap is not counted. - Without it, the cache stopped short of all experts, or was refused outright after a large file filled the page cache (45.29 GiB wanted, 44.64 GiB "free", 118 GiB available). - `hip_compat`: `cudaDevAttrIntegrated` → `hipDeviceAttributeIntegrated`. **3. setup** - A unified memory GPU is detected from the driver (`CU_DEVICE_ATTRIBUTE_INTEGRATED` via libcuda, matched to nvidia-smi by PCI address; `memory.total [N/A]` only as a fallback) and gets the system's RAM as its memory. - aarch64: no ready-made engine (the release is x86-64), so it compiles with `GGML_CPU_ARM_ARCH` from `/proc/cpuinfo` (engine and strata-vision); the CUDA repository key comes from `sbsa`; `cpu_info()` handles the absence of an x86 `flags` line; no AVX2 requirement. - Unified memory: `--mmap-experts` (an expert arena in RAM would be a second copy of the cache in the same memory) and no KV streaming. Plus `docs/DGX_SPARK.md` and a README pointer next to AMD HIP. ### Tested on a DGX Spark - `./setup.sh --model IQ2_XS --gguf-dir ... --vision no --yes`: detects the GB10, compiles, packs, writes the start script. - Correctness: coding, factual and arithmetic prompts answered correctly at 66-77 tok/s. - The CPU path: with `--expert-cache 2048` the CPU computed ~22 of ~29 experts per layer, Q2_0 rows included. Output was identical, token for token, to the all-GPU run (18.8 vs 81.9 tok/s). - `ctest`: 51 of 52 pass. `ple_parity` needs the Q2_0 model's second shard, which was not present. `expert_multi_test` (the AVX-512 Q2_0 kernel) is now registered on x86 only. - `tools/test_setup_*.py`: pass, except `test_setup_choices`, which fails the same way on unmodified `main` on this machine (`find_nvcc`). `test_setup_unsloth` simulates an x86 PC and now pins `setup.ARM` to `False`, as it already pins the CPU and GPU. ### Not covered - Models other than IQ2_XS were not run. All of their expert formats (IQ1_M, IQ2_XXS/XS/S, IQ3_XXS/S, IQ4_XS/NL, Q2_0; Q4_K/Q5_K/Q5_1/Q8_0 for the manual imports) have a ggml NEON path. - Images (`--vision`): built with the ARM flags, not tested. - The HIP build was not compiled (no ROCm here); the one HIP-side change is the attribute mapping above. - Strata's x86 multi-token CPU kernels have no NEON version. On a Spark the CPU computes no experts, so this only matters with a reduced cache. I'm happy to port them as a follow-up if you want ARM CPUs to carry experts. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.