v0.1.41
IA 125B na sua mesa
NVIDIA ou AMD · 12 GB+ VRAM · Windows e Linux

v0.1.41
Strata v0.1.41
Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group per GPU, and short prompt chunks on one NVIDIA GPU get help fro…
Para ler
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · anteontem · Benchmarks
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a rack. https://t.co/…
@arnabkarmkr · anteontem · Benchmarks
- Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090
@ItsmeAjayKV · anteontem · Benchmarks
- Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers nearly the same performa…
@thetechnotice · anteontem · Benchmarks

Nada sai do seu PC
Depois do setup, modelos e chats ficam neste PC.

60+ tok/s em GPUs RTX 5070
Na tabela: tokens por dia, energia e ponto de equilíbrio.

API compatível OpenAI
Baixe o motor e aponte o Cursor ou o Codex para 127.0.0.1:8080.
Deixe a IA instalar
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
Pulse
@coldniko
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
- Only loses to @ciruai engine on Hermes benchmarks. Kudos 🫡
- Strata wins by a tight margin against all other Strix Halo inference engines in it’s first try release. Did not expect that! 🥳 https://t.co…