v0.1.41
桌面上的 125B 级 AI
NVIDIA 或 AMD · 12GB+ 显存 · Windows 与 Linux · 社区 MIT 开源引擎。

v0.1.41
Strata v0.1.41
Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group per GPU, and short prompt chunks on one NVIDIA GPU get help fro…
现在值得看
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
- @servasyy_ai Amazing work from you! I love the community and all the help I get with great PR's and benchmarks you are doing. I'll be looking forward to your fu…
@coldniko · 前天 · 基准
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · 前天 · 基准
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a rack. https://t.co/…
@arnabkarmkr · 前天 · 基准

数据不出本机
安装之后,模型和对话都留在这台电脑上。

RTX 5070 级别可达 60+ tok/s
在基准页比较日吞吐、电费和回本天数。

OpenAI 兼容 API,对接 Cursor、Claude Code、Codex
下载引擎后,把 Cursor 或 Codex 接到 127.0.0.1:8080。
让 AI 帮你安装
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
作者动态
@coldniko 最新:Strix Halo、本地评测、Qwen3.8-Flash-Next。
- Only loses to @ciruai engine on Hermes benchmarks. Kudos 🫡
- Strata wins by a tight margin against all other Strix Halo inference engines in it’s first try release. Did not expect that! 🥳 https://t.co…