Pull requests / #1621
#1621 Add Jetson Orin ARM64 support and shared-memory budgeting
open · @calvinvette · 0 comments · View on GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Description
ARM64 hosts currently encounter x86-specific build and CPU assumptions, and Jetson Orin needs cache sizing that accounts for CPU and GPU sharing physical RAM. This adds native ARM64 build/setup support, portable CPU paths, and integrated-device memory handling, validated on a Jetson AGX Orin Developer Kit 32 GB with JetPack 6.2, CUDA 12.6 and SM87. Related: #1306. This validates Orin; GB10 hardware was not tested. ## Changes - Select ARM64 CPU sources and guard x86 intrinsics; retain the existing x86 paths and activation contracts. - Discover ARM CUDA devices through the driver API and preserve JetPack CUDA libraries during local installation. - Budget integrated-device expert caches against physical RAM, leaving six GiB headroom and a separate three GiB late-workspace reserve under automatic sizing; retain explicit user overrides. - Handle mapped/pinned memory limitations, add portable tests and platform probes, and fix the position-zero single-token execution sentinel. - Include reproducible validation and same-day context comparison evidence under `bench/results/2026-10-07-jetson-orin` and `bench/results/2026-10-08-jetson-orin`. ## Validation - Full Release CUDA SM87 build on upstream 0.1.41; 102 eligible CTest cases passed across the initial suite and targeted rerun. Three legacy fixture-dependent cases are excluded. - Real Coder IQ1_M native expert checks across all 48 layers and seven real-model API checks passed, including streaming and cancellation recovery. - Jetson setup: 15 cases; desktop setup: 111 cases; x86 CPU/router cross-build passed. - Full Release HIP gfx1100 build with MMQ and tests enabled, and full Release SYCL SPIR-V build with parity targets enabled, passed on `maestro1` (Celeron N3450, x86-64 Ubuntu 22.04). Toolchains: AMD TheRock `7.10.0a20251120` / Clang 22 and Intel DPC++ 2026.1.1 / oneMKL 2026.1.0. Both binaries loaded and printed help. The HIP shared-memory budget test passed; three AVX-512-dependent CPU checks skipped on this host. - Same-day original-port versus rebased-port sweep: all 18 cells at context capacities 1K–256K preserved six GiB physical headroom. 156 matched throughput requests, nine near-limit requests and nine recovery checks completed. - At 256K, the rebased build read 261,936 prompt tokens in 4,480.2 seconds and generated 64 tokens at 6.6 tokens/s, with at least 9.33 GiB available physical RAM. Swap was in use; clocks were not locked. Automatic cache allocations differed, so these are configuration observations rather than isolated kernel comparisons. Detailed methods, raw outputs and limitations: [validation report](https://github.com/calvinvette/strata-inference-orin/blob/feat/jetson-orin-arm64/bench/results/2026-10-08-jetson-orin/README.md). Backend compiler versions, complete build logs, commands, library loading checks and setup fixture exclusions: [maestro1 build report](https://github.com/calvinvette/strata-inference-orin/blob/feat/jetson-orin-arm64/bench/results/2026-10-08-maestro1-builds/README.md). These checks required no engine source changes. ## Remaining validation HIP/SYCL GPU tests and model execution, desktop GPU runtime, Windows, other models and GB10 were not tested. The x86 builds check compilation and linking; they do not validate ARM HIP/SYCL hosts or extend the Orin performance results to other GPUs. Desktop inference output byte identity remains unverified. Long-context execution does not establish recall accuracy or numerical equivalence across the full window.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.