Pull requests / #409

#409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory

closed · @eelgaev · 0 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Description

## aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory

This PR makes Strata build and run on aarch64, and `./setup.sh` work end to end on an NVIDIA DGX Spark. It is experimental and community-tested, in the same spirit as the AMD HIP and sm_70 builds: **one DGX Spark, Qwen3.8-Flash-Next IQ2_XS installed by setup**. x86 builds are unchanged.

### Speed (IQ2_XS, setup's own `run-iq2_xs.sh`: 128K context, `--kv int8`)

All 24,576 experts fit in the GPU's expert cache (33.0 GiB), so the CPU computes none.

| Prompt | Context | Prefill | Decode | MTP acceptance |
| ---: | ---: | ---: | ---: | ---: |
| 2K | 0 | 928 tok/s | 58.7 tok/s | 78% |
| 8K | 0 | 1,356 tok/s | 62.0 tok/s | 81% |
| 2K | 16K | 1,421 tok/s | 60.5 tok/s | 83% |
| 8K | 16K | 1,450 tok/s | 58.2 tok/s | 79% |
| 2K | 64K | 1,505 tok/s | 58.9 tok/s | 80% |
| 8K | 64K | 1,515 tok/s | 54.9 tok/s | 75% |

llama-benchy 0.4.0, 3 runs after a warm-up, `--no-cache`, one request at a time, 128 tokens decoded per request. Prefill and decode are from the server log (`strata serve: prompt ...`). benchy's `pp` column cannot time Strata's server, because the first stream chunk is sent before the prompt is read; details are in `docs/DGX_SPARK.md`.

### What changes

The change splits naturally into three parts. Happy to send them as separate PRs if you prefer.

**1. Build and engine on aarch64**
- `src/kernels/cpu/portable.cpp` replaces the AVX-only kernel files on non-x86 (CMake chooses by `CMAKE_SYSTEM_PROCESSOR`). It contains the activation quantizer (expert.cpp's scalar loop) and the BF16 router dot (NEON). The AVX multi-token kernels report themselves unsupported, so native experts use ggml-cpu's NEON `vec_dot`.
- Q2_0 expert rows (the down projections of every setup model) go through ggml-cpu on non-x86 (`q2_own_kernels()`). x86 keeps `q2_rows_any`.
- `include/strata/platform/cpu_relax.hpp`: `_mm_pause` / `_mm_sfence` on x86 (as before), `yield` / `dmb oshst` on aarch64.
- `-ffp-contract=off` for C/C++ on non-x86 GCC/Clang, as an x86 build without `-mfma` effectively has, so host roundings match x86 and ggml. (`quantize_act_parity`'s Q8_K case failed without it.)
- `GGML_CPU_ARM_ARCH` turns `GGML_NATIVE` off: gcc's `-mcpu=native` finds no features on GB10's Cortex-X925/A725, which leaves ggml-cpu without dotprod/i8mm.

**2. Expert cache sizing on unified memory**
- `device_free_bytes()` is used by both the cache sizing and `ExpertCache::open`'s check. When the driver reports that the GPU shares host memory (`cudaDevAttrIntegrated`), it counts `MemAvailable` less 6 GiB (`STRATA_UMA_HEADROOM_GIB`) as free. `cudaMemGetInfo` counts only free RAM, and with `--mmap-experts` the page cache is mostly the GGUF's own pages. NVIDIA's DGX Spark Porting Guide (§5.5) recommends the same; swap is not counted.
- Without it, the cache stopped short of all experts, or was refused outright after a large file filled the page cache (45.29 GiB wanted, 44.64 GiB "free", 118 GiB available).
- `hip_compat`: `cudaDevAttrIntegrated` → `hipDeviceAttributeIntegrated`.

**3. setup**
- A unified memory GPU is detected from the driver (`CU_DEVICE_ATTRIBUTE_INTEGRATED` via libcuda, matched to nvidia-smi by PCI address; `memory.total [N/A]` only as a fallback) and gets the system's RAM as its memory.
- aarch64: no ready-made engine (the release is x86-64), so it compiles with `GGML_CPU_ARM_ARCH` from `/proc/cpuinfo` (engine and strata-vision); the CUDA repository key comes from `sbsa`; `cpu_info()` handles the absence of an x86 `flags` line; no AVX2 requirement.
- Unified memory: `--mmap-experts` (an expert arena in RAM would be a second copy of the cache in the same memory) and no KV streaming.

Plus `docs/DGX_SPARK.md` and a README pointer next to AMD HIP.

### Tested on a DGX Spark

- `./setup.sh --model IQ2_XS --gguf-dir ... --vision no --yes`: detects the GB10, compiles, packs, writes the start script.
- Correctness: coding, factual and arithmetic prompts answered correctly at 66-77 tok/s.
- The CPU path: with `--expert-cache 2048` the CPU computed ~22 of ~29 experts per layer, Q2_0 rows included. Output was identical, token for token, to the all-GPU run (18.8 vs 81.9 tok/s).
- `ctest`: 51 of 52 pass. `ple_parity` needs the Q2_0 model's second shard, which was not present. `expert_multi_test` (the AVX-512 Q2_0 kernel) is now registered on x86 only.
- `tools/test_setup_*.py`: pass, except `test_setup_choices`, which fails the same way on unmodified `main` on this machine (`find_nvcc`). `test_setup_unsloth` simulates an x86 PC and now pins `setup.ARM` to `False`, as it already pins the CPU and GPU.

### Not covered

- Models other than IQ2_XS were not run. All of their expert formats (IQ1_M, IQ2_XXS/XS/S, IQ3_XXS/S, IQ4_XS/NL, Q2_0; Q4_K/Q5_K/Q5_1/Q8_0 for the manual imports) have a ggml NEON path.
- Images (`--vision`): built with the ARM flags, not tested.
- The HIP build was not compiled (no ROCm here); the one HIP-side change is the attribute mapping above.
- Strata's x86 multi-token CPU kernels have no NEON version. On a Spark the CPU computes no experts, so this only matters with a reduced cache. I'm happy to port them as a follow-up if you want ARM CPUs to carry experts.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.