Pull requests / #1449
#1449 Add Metal backend support for Strata on macOS
open · @dennis-akimov · 0 comments · View on GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation
Description
## Title Experimental macOS / Apple Silicon support (Metal engine) ## Summary Strata now runs on Apple Silicon Macs (M1 or newer, 64 GB+). A new engine, `strata-metal`, runs the model on llama.cpp's Metal backend and speaks the CUDA engine's line protocol, so the server, web app and APIs (OpenAI, Anthropic, Responses, MCP) work unchanged. Setup detects a Mac and builds it by itself: `make check`, `make pull MODEL=Q2_0`, `make run`. Experimental: tested on one MacBook Pro M5 Max (128 GB, macOS 26.4) with Q2_0. Measured on that Mac (Q2_0, 32K context, through the OpenAI API, greedy, thinking off): - Writing the answer: 13–17 tok/s with other programs running, up to 48.3 tok/s on a quieter machine. With `--mtp on` (opt-in): up to 21.6 tok/s; MTP alone gives 1.23–1.50x on short chats. - Reading the prompt: 225–272 tok/s, up to ~900 tok/s on a 3.3K-token prompt in a quieter run. - A follow-up message in a long conversation starts answering in 0.3 s. - mlx-lm on the same weights (converted to MLX) for comparison: up to 41.5 tok/s, with ~81 GB in memory vs ~35 GB. The README's one-shot voxel pagoda garden prompt, with a 4,096-token thinking budget, produced a complete HTML page in 443 s (~3.7 min of capped thinking, then ~16K characters of code at ~20 tok/s). Without the budget, greedy thinking ran for 90K tokens and never answered. ## What changed - `metal/strata_metal.cpp`: the Metal engine: conversation reuse through recurrent-state checkpoints, pictures with M-RoPE positions, batch slots, MTP drafting. - `metal/CMakeLists.txt`, `CMakeLists.txt`: `-DSTRATA_ENABLE_METAL=ON`, llama.cpp at a pinned commit; `metal/patches/`: two optional Metal kernel patches; `metal/bench/ab.py`: A/B test of two engine builds. - `metal/setup_mac.py`, `setup.py`, `setup.sh`: setup's Mac steps (Metal GPU check, local build, config conversion); `make check` on a Mac uses a 64 GB floor and marks untested sizes. - `metal/mtp_gguf.py`: builds the MTP draft head for llama.cpp, checked by SHA-256. - `serve/telemetry.py`, `serve/web/app.js`: Mac GPU load, memory, power and chip temperature in the Monitor, without root; PCIe shows "n/a" for an integrated GPU. `serve/server.py`: RFC 6750 `WWW-Authenticate` on 401. - `Makefile`, `tools/chat.py`, `tools/set_context.py`: `make run/start/stop/status/chat/pull/check/test`, `make run CONTEXT=131072`, `make chat REASONING_BUDGET=N`; works with macOS's make 3.81. - `docs/MACOS.md`: quick start, limits and warnings, measurements, MLX comparison; a line in each of the 7 READMEs. - Tests: engine protocol, setup conversion, Mac telemetry (with a live Apple Silicon test), Monitor tiles (Node), `make chat` and the context option. Three tests that failed only on macOS now pass; `make test` passes on the Mac. ## Extra Notes - Only Q2_0 was tested on a Mac; `make check` marks the other sizes "untested on a Mac". - Not on a Mac yet: contexts past 262,144 tokens (rope scaling), the PC expert cache, CPU experts. - Power and temperature use private macOS interfaces; if a macOS update changes them, the tiles show "–". - Decode speed depends a lot on background load (a VM plus an Xcode build halved it) and on laptop power settings.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.