Pull requests / #901

#901 Add ARM64 NVIDIA GB10 build and native-pack support

closed · @githubhjs · 0 comentarios · En GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quants

Descripción

## ARM64 / NVIDIA GB10 support

This PR adds a tested ARM64 build path for NVIDIA GB10 (Grace Blackwell, compute capability 12.1).

### Changes

- Detect GB10 unified memory when `nvidia-smi` reports `[N/A]` for VRAM.
- Build CUDA kernels for `sm_121`.
- Add scalar ARM64 Q2_0 expert and BF16 router fallbacks.
- Make CPU pause/fence paths portable to ARM64.
- Enable ARM64 native packs through ggml CPU, while skipping x86-only AVX multi-token kernels.
- Allow setup and the engine to recognize ARM64 instead of applying the x86 SSE4.2 floor.

### Hardware validation

On NVIDIA GB10, 121.6 GiB unified GPU memory, CUDA 13.0:

- Full engine build completed with `STRATA_NATIVE_EXPERTS=ON` and CUDA architecture 121.
- `strata-device --list-devices` detected GB10, compute capability 12.1.
- Qwen3.8-Flash-Next GSQ-RCO Q2_0 model and MTP draft layer loaded successfully.
- OpenAI-compatible generation returned `4` for `What is 2 + 2?`.
- Measured prompt throughput: 73.6 tok/s.
- Measured generation throughput: 48.7 tok/s.
- MTP draft acceptance: 20 / 25 proposed tokens.

The ARM64 CPU expert path is scalar and intended as a correctness fallback; the tested GB10 run uses the CUDA path for the main model work.

En el sitio

Enlaces a install, modelos, releases.