v0.1.41
125B-class AI on your desk
NVIDIA or AMD · 12 GB+ VRAM · Windows & Linux · MIT licensed engine from the community.

v0.1.41
Strata v0.1.41
Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group per GPU, and short prompt chunks on one NVIDIA GPU get help fro…
Worth reading now
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
@coldniko · 2 days ago · Releases · Benchmarks
- @servasyy_ai Amazing work from you! I love the community and all the help I get with great PR's and benchmarks you are doing. I'll be looking forward to your fu…
@coldniko · 2 days ago · Benchmarks
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · 2 days ago · Benchmarks
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a rack. https://t.co/…
@arnabkarmkr · 2 days ago · Benchmarks

Nothing leaves your PC
Setup keeps the models and the chats on this PC.

60+ tokens/s on RTX 5070 class GPUs
Compare daily tokens, power cost, and break-even on the benchmark board.

OpenAI-compatible API for Cursor, Claude Code, Codex
Download the engine, then point Cursor or Codex at 127.0.0.1:8080.
Let your AI install Strata
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
Creator pulse
Latest from @coldniko — Strix Halo, local benchmarks, Qwen3.8-Flash-Next.
- Only loses to @ciruai engine on Hermes benchmarks. Kudos 🫡
- Strata wins by a tight margin against all other Strix Halo inference engines in it’s first try release. Did not expect that! 🥳 https://t.co…