Pull requests / #1449
#1449 Add Metal backend support for Strata on macOS
open · @dennis-akimov · 0 comentarios · En GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentation
Descripción
## Title Experimental macOS / Apple Silicon support (Metal engine) ## Summary Strata now runs on Apple Silicon Macs (M1 or newer, 64 GB+). A new engine, `strata-metal`, runs the model on llama.cpp's Metal backend and speaks the CUDA engine's line protocol, so the server, web app and APIs (OpenAI, Anthropic, Responses, MCP) work unchanged. Setup detects a Mac and builds it by itself: `make check`, `make pull MODEL=Q2_0`, `make run`. Experimental: tested on one MacBook Pro M5 Max (128 GB, macOS 26.4) with Q2_0. Measured on that Mac (Q2_0, 32K context, through the OpenAI API, greedy, thinking off): - Writing the answer: 13–17 tok/s with other programs running, up to 48.3 tok/s on a quieter machine. With `--mtp on` (opt-in): up to 21.6 tok/s; MTP alone gives 1.23–1.50x on short chats. - Reading the prompt: 225–272 tok/s, up to ~900 tok/s on a 3.3K-token prompt in a quieter run. - A follow-up message in a long conversation starts answering in 0.3 s. - mlx-lm on the same weights (converted to MLX) for comparison: up to 41.5 tok/s, with ~81 GB in memory vs ~35 GB. The README's one-shot voxel pagoda garden prompt, with a 4,096-token thinking budget, produced a complete HTML page in 443 s (~3.7 min of capped thinking, then ~16K characters of code at ~20 tok/s). Without the budget, greedy thinking ran for 90K tokens and never answered. ## What changed - `metal/strata_metal.cpp`: the Metal engine: conversation reuse through recurrent-state checkpoints, pictures with M-RoPE positions, batch slots, MTP drafting. - `metal/CMakeLists.txt`, `CMakeLists.txt`: `-DSTRATA_ENABLE_METAL=ON`, llama.cpp at a pinned commit; `metal/patches/`: two optional Metal kernel patches; `metal/bench/ab.py`: A/B test of two engine builds. - `metal/setup_mac.py`, `setup.py`, `setup.sh`: setup's Mac steps (Metal GPU check, local build, config conversion); `make check` on a Mac uses a 64 GB floor and marks untested sizes. - `metal/mtp_gguf.py`: builds the MTP draft head for llama.cpp, checked by SHA-256. - `serve/telemetry.py`, `serve/web/app.js`: Mac GPU load, memory, power and chip temperature in the Monitor, without root; PCIe shows "n/a" for an integrated GPU. `serve/server.py`: RFC 6750 `WWW-Authenticate` on 401. - `Makefile`, `tools/chat.py`, `tools/set_context.py`: `make run/start/stop/status/chat/pull/check/test`, `make run CONTEXT=131072`, `make chat REASONING_BUDGET=N`; works with macOS's make 3.81. - `docs/MACOS.md`: quick start, limits and warnings, measurements, MLX comparison; a line in each of the 7 READMEs. - Tests: engine protocol, setup conversion, Mac telemetry (with a live Apple Silicon test), Monitor tiles (Node), `make chat` and the context option. Three tests that failed only on macOS now pass; `make test` passes on the Mac. ## Extra Notes - Only Q2_0 was tested on a Mac; `make check` marks the other sizes "untested on a Mac". - Not on a Mac yet: contexts past 262,144 tokens (rope scaling), the PC expert cache, CPU experts. - Power and temperature use private macOS interfaces; if a macOS update changes them, the tiles show "–". - Decode speed depends a lot on background load (a VM plus an Xcode build halved it) and on laptop power settings.
En el sitio
Enlaces a install, modelos, releases.