v0.1.41
125B-KI auf dem Schreibtisch
NVIDIA oder AMD · 12 GB+ VRAM · Windows & Linux

v0.1.41
Strata v0.1.41
Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group per GPU, and short prompt chunks on one NVIDIA GPU get help fro…
Jetzt lesen
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · vorgestern · Benchmarks
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a rack. https://t.co/…
@arnabkarmkr · vorgestern · Benchmarks
- Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090
@ItsmeAjayKV · vorgestern · Benchmarks
- Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers nearly the same performa…
@thetechnotice · vorgestern · Benchmarks

Nichts verlässt deinen PC
Nach dem Setup bleiben Modelle und Chats auf diesem PC.

60+ tok/s auf RTX-5070-Klasse
Auf der Benchmark-Seite: Tokens pro Tag, Strom und Amortisation.

OpenAI-kompatible API
Engine laden und Cursor oder Codex auf 127.0.0.1:8080 legen.
KI installiert Strata
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.
Autor-Feed
Neues von @coldniko
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
- Only loses to @ciruai engine on Hermes benchmarks. Kudos 🫡
- Strata wins by a tight margin against all other Strix Halo inference engines in it’s first try release. Did not expect that! 🥳 https://t.co…