Pull requests / #1449

#1449 Add Metal backend support for Strata on macOS

open · @dennis-akimov · 0 comentarios · En GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation

Descripción

## Title
Experimental macOS / Apple Silicon support (Metal engine)

## Summary
Strata now runs on Apple Silicon Macs (M1 or newer, 64 GB+). A new engine, `strata-metal`, runs the model on
llama.cpp's Metal backend and speaks the CUDA engine's line protocol, so the server, web app and APIs (OpenAI,
Anthropic, Responses, MCP) work unchanged. Setup detects a Mac and builds it by itself: `make check`,
`make pull MODEL=Q2_0`, `make run`. Experimental: tested on one MacBook Pro M5 Max (128 GB, macOS 26.4) with Q2_0.

Measured on that Mac (Q2_0, 32K context, through the OpenAI API, greedy, thinking off):
- Writing the answer: 13–17 tok/s with other programs running, up to 48.3 tok/s on a quieter machine.
  With `--mtp on` (opt-in): up to 21.6 tok/s; MTP alone gives 1.23–1.50x on short chats.
- Reading the prompt: 225–272 tok/s, up to ~900 tok/s on a 3.3K-token prompt in a quieter run.
- A follow-up message in a long conversation starts answering in 0.3 s.
- mlx-lm on the same weights (converted to MLX) for comparison: up to 41.5 tok/s, with ~81 GB in memory vs ~35 GB.

The README's one-shot voxel pagoda garden prompt, with a 4,096-token thinking budget, produced a complete HTML
page in 443 s (~3.7 min of capped thinking, then ~16K characters of code at ~20 tok/s). Without the budget,
greedy thinking ran for 90K tokens and never answered.

## What changed
- `metal/strata_metal.cpp`: the Metal engine: conversation reuse through recurrent-state checkpoints, pictures
  with M-RoPE positions, batch slots, MTP drafting.
- `metal/CMakeLists.txt`, `CMakeLists.txt`: `-DSTRATA_ENABLE_METAL=ON`, llama.cpp at a pinned commit;
  `metal/patches/`: two optional Metal kernel patches; `metal/bench/ab.py`: A/B test of two engine builds.
- `metal/setup_mac.py`, `setup.py`, `setup.sh`: setup's Mac steps (Metal GPU check, local build, config
  conversion); `make check` on a Mac uses a 64 GB floor and marks untested sizes.
- `metal/mtp_gguf.py`: builds the MTP draft head for llama.cpp, checked by SHA-256.
- `serve/telemetry.py`, `serve/web/app.js`: Mac GPU load, memory, power and chip temperature in the Monitor,
  without root; PCIe shows "n/a" for an integrated GPU. `serve/server.py`: RFC 6750 `WWW-Authenticate` on 401.
- `Makefile`, `tools/chat.py`, `tools/set_context.py`: `make run/start/stop/status/chat/pull/check/test`,
  `make run CONTEXT=131072`, `make chat REASONING_BUDGET=N`; works with macOS's make 3.81.
- `docs/MACOS.md`: quick start, limits and warnings, measurements, MLX comparison; a line in each of the 7 READMEs.
- Tests: engine protocol, setup conversion, Mac telemetry (with a live Apple Silicon test), Monitor tiles (Node),
  `make chat` and the context option. Three tests that failed only on macOS now pass; `make test` passes on the Mac.

## Extra Notes
- Only Q2_0 was tested on a Mac; `make check` marks the other sizes "untested on a Mac".
- Not on a Mac yet: contexts past 262,144 tokens (rope scaling), the PC expert cache, CPU experts.
- Power and temperature use private macOS interfaces; if a macOS update changes them, the tiles show "–".
- Decode speed depends a lot on background load (a VM plus an Xcode build halved it) and on laptop power settings.

En el sitio

Enlaces a install, modelos, releases.