Pull requests / #1621

#1621 Add Jetson Orin ARM64 support and shared-memory budgeting

open · @calvinvette · 0 Kommentare · Auf GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows

Beschreibung

ARM64 hosts currently encounter x86-specific build and CPU assumptions, and Jetson Orin needs cache sizing that accounts for CPU and GPU sharing physical RAM. This adds native ARM64 build/setup support, portable CPU paths, and integrated-device memory handling, validated on a Jetson AGX Orin Developer Kit 32 GB with JetPack 6.2, CUDA 12.6 and SM87.

Related: #1306. This validates Orin; GB10 hardware was not tested.

## Changes

- Select ARM64 CPU sources and guard x86 intrinsics; retain the existing x86 paths and activation contracts.
- Discover ARM CUDA devices through the driver API and preserve JetPack CUDA libraries during local installation.
- Budget integrated-device expert caches against physical RAM, leaving six GiB headroom and a separate three GiB late-workspace reserve under automatic sizing; retain explicit user overrides.
- Handle mapped/pinned memory limitations, add portable tests and platform probes, and fix the position-zero single-token execution sentinel.
- Include reproducible validation and same-day context comparison evidence under `bench/results/2026-10-07-jetson-orin` and `bench/results/2026-10-08-jetson-orin`.

## Validation

- Full Release CUDA SM87 build on upstream 0.1.41; 102 eligible CTest cases passed across the initial suite and targeted rerun. Three legacy fixture-dependent cases are excluded.
- Real Coder IQ1_M native expert checks across all 48 layers and seven real-model API checks passed, including streaming and cancellation recovery.
- Jetson setup: 15 cases; desktop setup: 111 cases; x86 CPU/router cross-build passed.
- Full Release HIP gfx1100 build with MMQ and tests enabled, and full Release SYCL SPIR-V build with parity targets enabled, passed on `maestro1` (Celeron N3450, x86-64 Ubuntu 22.04). Toolchains: AMD TheRock `7.10.0a20251120` / Clang 22 and Intel DPC++ 2026.1.1 / oneMKL 2026.1.0. Both binaries loaded and printed help. The HIP shared-memory budget test passed; three AVX-512-dependent CPU checks skipped on this host.
- Same-day original-port versus rebased-port sweep: all 18 cells at context capacities 1K–256K preserved six GiB physical headroom. 156 matched throughput requests, nine near-limit requests and nine recovery checks completed.
- At 256K, the rebased build read 261,936 prompt tokens in 4,480.2 seconds and generated 64 tokens at 6.6 tokens/s, with at least 9.33 GiB available physical RAM. Swap was in use; clocks were not locked. Automatic cache allocations differed, so these are configuration observations rather than isolated kernel comparisons.

Detailed methods, raw outputs and limitations: [validation report](https://github.com/calvinvette/strata-inference-orin/blob/feat/jetson-orin-arm64/bench/results/2026-10-08-jetson-orin/README.md).

Backend compiler versions, complete build logs, commands, library loading checks and setup fixture exclusions: [maestro1 build report](https://github.com/calvinvette/strata-inference-orin/blob/feat/jetson-orin-arm64/bench/results/2026-10-08-maestro1-builds/README.md). These checks required no engine source changes.

## Remaining validation

HIP/SYCL GPU tests and model execution, desktop GPU runtime, Windows, other models and GB10 were not tested. The x86 builds check compilation and linking; they do not validate ARM HIP/SYCL hosts or extend the Orin performance results to other GPUs. Desktop inference output byte identity remains unverified. Long-context execution does not establish recall accuracy or numerical equivalence across the full window.

Mehr auf der Site

Links zu Install, Modellen, Releases.