Pull requests / #901
#901 Add ARM64 NVIDIA GB10 build and native-pack support
closed · @githubhjs · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quants
Beschreibung
## ARM64 / NVIDIA GB10 support This PR adds a tested ARM64 build path for NVIDIA GB10 (Grace Blackwell, compute capability 12.1). ### Changes - Detect GB10 unified memory when `nvidia-smi` reports `[N/A]` for VRAM. - Build CUDA kernels for `sm_121`. - Add scalar ARM64 Q2_0 expert and BF16 router fallbacks. - Make CPU pause/fence paths portable to ARM64. - Enable ARM64 native packs through ggml CPU, while skipping x86-only AVX multi-token kernels. - Allow setup and the engine to recognize ARM64 instead of applying the x86 SSE4.2 floor. ### Hardware validation On NVIDIA GB10, 121.6 GiB unified GPU memory, CUDA 13.0: - Full engine build completed with `STRATA_NATIVE_EXPERTS=ON` and CUDA architecture 121. - `strata-device --list-devices` detected GB10, compute capability 12.1. - Qwen3.8-Flash-Next GSQ-RCO Q2_0 model and MTP draft layer loaded successfully. - OpenAI-compatible generation returned `4` for `What is 2 + 2?`. - Measured prompt throughput: 73.6 tok/s. - Measured generation throughput: 48.7 tok/s. - MTP draft acceptance: 20 / 25 proposed tokens. The ARM64 CPU expert path is scalar and intended as a correctness fallback; the tested GB10 run uses the CUDA path for the main model work.
Mehr auf der Site
Links zu Install, Modellen, Releases.