Intelligence pulse

X feed: local LLM, Strata, GPU prices.
Synced 2026-10-08 17:27 UTC · Neuester Post vorgestern · cache · Verify quotes.
Jetzt lesen
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · vorgestern · Benchmarks
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a rack. https://t.co/…
@arnabkarmkr · vorgestern · Benchmarks
- Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090
@ItsmeAjayKV · vorgestern · Benchmarks
- Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers nearly the same performa…
@thetechnotice · vorgestern · Benchmarks
- Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strata (MTP on). Cache re…
@AgentWorkflowLa · vorgestern · Benchmarks
- another few ns here
@CCDavidCode · vorgestern · Benchmarks
Latest on X
Posts from tracked accounts and Strata-related searches (RSSHub / X API).
- @leypascua @HuggingModels X reports: 150-180 tok/s on RTX 5090 (Strata), 35-45 on 20-series for the Coder GGUF; 44 tps claimed on 32GB+ syst…
@grok · vorgestern · X · benchmark, models, strata
@leypascua @HuggingModels X reports: 150-180 tok/s on RTX 5090 (Strata), 35-45 on 20-series for the Coder GGUF; 44 tps claimed on 32GB+ systems. Reddit/HF: 29.6 GB resident works with SSD n-gram mmap. Coding benches hold…
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · vorgestern · X · benchmark, models, strata
Old Ethereum mining GPUs got a new job: running a 125B model. Qwen3.8-Flash-Next + Strata, 16/24 GB GPUs, 64 GB RAM. Up to 138.7 tok/s. Real benchmarks, first-token latency, and cached turns. https://t.co/gFM9KnQeVq #Loc…
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a…
@arnabkarmkr · vorgestern · X · benchmark, models, strata
- @Androthon @Oluwaphilemon1 Qwen3.8-Flash-Next (125B MoE, ~6B active + 51B n-gram). Inference server is Strata. Needs 12GB+ VRAM and 64GB+ sy…
@grok · vorgestern · X · benchmark, models, strata
@Androthon @Oluwaphilemon1 Qwen3.8-Flash-Next (125B MoE, ~6B active + 51B n-gram). Inference server is Strata. Needs 12GB+ VRAM and 64GB+ system RAM for solid performance (32GB RAM minimum for lighter quants like Coder).…
- @Androthon @Oluwaphilemon1 It's just Strata. Strata is the complete open-source inference engine for Qwen3.8-Flash-Next. It handles expert c…
@grok · vorgestern · X · models, strata
@Androthon @Oluwaphilemon1 It's just Strata. Strata is the complete open-source inference engine for Qwen3.8-Flash-Next. It handles expert caching (VRAM/RAM/SSD), memory tiering, and speculative decoding on its own, then…
- Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090
@ItsmeAjayKV · vorgestern · X · models, strata, gpu
Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090 Here is a jungle fps theme game one-shot generated by IQ3_S System: - RTX 3090 - 64GB RAM - Strata Inference engine IQ3_S does produce better graphics …
- Single DGX Spark owners have a new model worth paying attention to.
@Oluwaphilemon1 · vorgestern · X · models, deploy
Single DGX Spark owners have a new model worth paying attention to. Qwen3.8-Flash-Next NVFP4 can now be served on one Spark with a setup that is surprisingly capable for a single 128GB unified-memory machine. The headlin…
- MLX-Serve 26.9.5 is starting to look less like a local model runner and more like an actual AI serving stack for Apple Silicon.
@Oluwaphilemon1 · vorgestern · X · models
MLX-Serve 26.9.5 is starting to look less like a local model runner and more like an actual AI serving stack for Apple Silicon. The headline additions are Qwen-Image-2.1 and Prism Bonsai 2. But the concurrency work is ar…
- @ItsmeAjayKV @0xSero @coldniko came here looking for this those tok/s numbers look anemic compared to what I'm experiencing with Strata
@rskjr · vorgestern · X · benchmark, strata, comparison
- NVIDIA launched its high-end RTX 50 "Blackwell" graphics cards, with the RTX 5090 as the flagship model.
@investies · vorgestern · X · release, gpu, pricing
NVIDIA launched its high-end RTX 50 "Blackwell" graphics cards, with the RTX 5090 as the flagship model. The RTX 5090 boasts impressive specs and a price tag of $1,999. What's the first game you'd play with this powerhou…
- Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers near…
@thetechnotice · vorgestern · X · benchmark, gpu, pricing
Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers nearly the same performance for a fraction of the price. This is the card to get! #G…
- Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strat…
@AgentWorkflowLa · vorgestern · X · benchmark, models, strata
Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strata (MTP on). Cache reuse narrows the gap. Full benchmark: https://t.co/cpx6NOcU9i
- @LundukeJournal But it's a sad state if all the tokens are cloud based. Do at least some of the work on local LLMs. Strata is the best curre…
@Robert_of_Maine · vorgestern · X · strata
@LundukeJournal But it's a sad state if all the tokens are cloud based. Do at least some of the work on local LLMs. Strata is the best current inference engine for this if you've got at least 32GB of system ram.
- Prime Big Deal Days is live until 11:59pm PT tomorrow. New drops at midnight, 8am and 1pm Pacific.
@bottleneck_pc · vorgestern · X · gpu, pricing
Prime Big Deal Days is live until 11:59pm PT tomorrow. New drops at midnight, 8am and 1pm Pacific. The three lines to refresh for: RX 9070 XT under $700, Ryzen 7 9800X3D under $420, Corsair RM850e under $85. RAM and SSD …
- another few ns here
@CCDavidCode · vorgestern · X · benchmark, models, strata
another few ns here saved memory 8 times there. hmm dont need to wait on that Q8, plain decode llama.cpp: 22.0 tok/s Q8, MTP Early Q8 Strata prototype: 87.5 tok/s coding at 8K Strata now: 156.2 tok/s at 32K
- @_imdawon Strata is going to get faster
@CCDavidCode · vorgestern · X · benchmark, strata
- @coldniko Strix Halo 385 w/ 64gb memory getting 42 tp/s with strata IQ3_ XXS - on my llama.cpp build I was getting peak 27. Great job
@Prod_Philosophy · vorgestern · X · hardware, models, strata
- @IDFkushoverlord for dense models like qwen27B it was very slow for me even with all my efforts i could get preprocessing fast enough for my…
@l3d3ka · vorgestern · X · benchmark, models, strata
@IDFkushoverlord for dense models like qwen27B it was very slow for me even with all my efforts i could get preprocessing fast enough for my liking but since literally today i switched to strata inference engine and i’ve…
- 450+ tok/s of prefill on AMD Strix Halo.
@msala9 · vorgestern · X · hardware, benchmark, models
450+ tok/s of prefill on AMD Strix Halo. DS4 Halo accelerates DeepSeek V4 Flash on a single Radeon 8060S integrated GPU, in Strix Halo 128 GB. 454.59 tokens/s measured on a complete 4K prompt. Model: DeepSeek V4 Flash 07…
- @rskjr Hey Robert! I have seen your live stream and concerns in Strata and totally agree. Things are progressing a bit fast, just wanted to …
@coldniko · vorgestern · X · update, strata
- @JJismusic @MiaAI_lab Yes, that's an effective approach. EmbeddingGemma 2's 740M modular size loads easily alongside Qwen3.8-Flash-Next as a…
@grok · vorgestern · X · models
@JJismusic @MiaAI_lab Yes, that's an effective approach. EmbeddingGemma 2's 740M modular size loads easily alongside Qwen3.8-Flash-Next as a retrieval tool. Use it for unified search across text/code/images/video/audio, …
- @0xSero 5080 (16GB)
@hugovntr · vorgestern · X · benchmark, strata
- Strata is pushing Qwen3.8-Flash-Next to ~165 tok/s on my RTX 5090. 🤯
@XniX · vorgestern · X · benchmark, models, strata
Strata is pushing Qwen3.8-Flash-Next to ~165 tok/s on my RTX 5090. 🤯 524K context + Vision enabled, running a 125B-class MoE locally on a single consumer GPU. Seriously impressive. 🔥 #LocalLLM #Qwen #RTX5090
- Same PC, same model files, same prompts. Upstream Strata 0.1.40.1 vs strata-rdna4, median of 3 runs:
@vectorweft · vorgestern · X · benchmark, strata, comparison
Same PC, same model files, same prompts. Upstream Strata 0.1.40.1 vs strata-rdna4, median of 3 runs: 📥 Prompt 32K: 515 → 1,826 tok/s 📥 Prompt 128K: 493 → 1,415 tok/s 📤 Output 32K: 51.5 → 84.7 tok/s ⏱️ First token at 1…
- @coldniko's Strata runs a 125B MoE model on a desktop PC. I spent the last week tuning it for AMD RDNA4.
@vectorweft · vorgestern · X · benchmark, models, strata
@coldniko's Strata runs a 125B MoE model on a desktop PC. I spent the last week tuning it for AMD RDNA4. Qwen3.8-Flash-Next (Unsloth UD-Q4_K_XL) on 2x Radeon AI PRO R9700: ⚡ 1,826 tok/s prompt reading at 32K ⚡ 85-100 tok…
- @yume_arasaki Strata is an amazing project! But the benchmarks for the full precision model is NOT what you are getting with Strata. In my t…
@realjohnmock · vorgestern · X · benchmark, models, strata
@yume_arasaki Strata is an amazing project! But the benchmarks for the full precision model is NOT what you are getting with Strata. In my testing it's on par or slightly below with Qwen 3.8 27b at a similar tok/s rate.
- Built on Qwen3.8 architecture with MoE (Mixture of Experts) and Flash-Next efficiency. Abliterated to remove refusal directions. GGUF format…
@HuggingModels · vorgestern · X · models, comparison
Built on Qwen3.8 architecture with MoE (Mixture of Experts) and Flash-Next efficiency. Abliterated to remove refusal directions. GGUF format for easy local inference. Vision-language capable. Optimized for llama.cpp. A p…
- 125B MoE on a laptop: 55 → 105 tok/s in 3 days 🚀
@baltaci212 · vorgestern · X · benchmark, models, strata
125B MoE on a laptop: 55 → 105 tok/s in 3 days 🚀 RTX 5070 Ti Laptop (12 GB) + 64 GB RAM, Qwen3.8-Flash-Next Swift 1.5 IQ3_XXS on the Strata engine • 30-question set: 55 → 105 tok/s (30/30 correct) • Turkish answers: 94 …
- GLM 5.3 Flash on a single Tesla V100 32GB + 64GB of RAM
@PeasantSmith · vorgestern · X · benchmark, models, strata
GLM 5.3 Flash on a single Tesla V100 32GB + 64GB of RAM -18B Active parameters -25 tok/s peak decode -@UnslothAI IQ1_S (93GB file) -CPU I5-12600T +DDR5 5200 + Gen4 NVME -Custom strata The Volta GPU is probably the floor,…
- @ItsmeAjayKV They are preparing for enhanced n-grams most likely. Strata shows that it possible to run them faster and more efficient as wel…
@Bitcopath · vorgestern · X · benchmark, strata
@ItsmeAjayKV They are preparing for enhanced n-grams most likely. Strata shows that it possible to run them faster and more efficient as well. I'm telling you for some time, we'll able to use much larger models with our …
- We entered the age of big models at home!
@IronWolve · vorgestern · X · benchmark, models, strata
We entered the age of big models at home! Using strata and a tweaked next flash, it rocks. Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF about 150-200 tok/s on my 5090 about 35-45 tok/s on my 2070 Super on a laptop
- @0xSero Using strata and a tweaked next flash, it rocks.
@IronWolve · vorgestern · X · benchmark, models, strata
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
@coldniko · vorgestern · X · benchmark, hardware, release
Community benchmark attention on Ryzen AI Max / Strix Halo builds with Strata v0.1.40.
- Atomic Chat shipped a TensorRT integration into their engine last week. Qwen3.5-4B and Qwen3-8B through CUDA on a Blackwell RTX 5090: 3.43 s…
@0xPascual · vorgestern · X · models, release, gpu
Atomic Chat shipped a TensorRT integration into their engine last week. Qwen3.5-4B and Qwen3-8B through CUDA on a Blackwell RTX 5090: 3.43 seconds, against 10.49 seconds on llama.cpp. A 206% throughput gap on identical s…
Comparisons & benchmarks · Upgrades & migration
Comparisons & benchmarks
Strata vs llama.cpp, vllm, tok/s threads, and MoE comparisons.
- Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090
@ItsmeAjayKV · vorgestern · X · models, strata, gpu
Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090 Here is a jungle fps theme game one-shot generated by IQ3_S System: - RTX 3090 - 64GB RAM - Strata Inference engine IQ3_S does produce better graphics …
- @ItsmeAjayKV @0xSero @coldniko came here looking for this those tok/s numbers look anemic compared to what I'm experiencing with Strata
@rskjr · vorgestern · X · benchmark, strata, comparison
- Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strat…
@AgentWorkflowLa · vorgestern · X · benchmark, models, strata
Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strata (MTP on). Cache reuse narrows the gap. Full benchmark: https://t.co/cpx6NOcU9i
- another few ns here
@CCDavidCode · vorgestern · X · benchmark, models, strata
another few ns here saved memory 8 times there. hmm dont need to wait on that Q8, plain decode llama.cpp: 22.0 tok/s Q8, MTP Early Q8 Strata prototype: 87.5 tok/s coding at 8K Strata now: 156.2 tok/s at 32K
- @coldniko Strix Halo 385 w/ 64gb memory getting 42 tp/s with strata IQ3_ XXS - on my llama.cpp build I was getting peak 27. Great job
@Prod_Philosophy · vorgestern · X · hardware, models, strata
- Same PC, same model files, same prompts. Upstream Strata 0.1.40.1 vs strata-rdna4, median of 3 runs:
@vectorweft · vorgestern · X · benchmark, strata, comparison
Same PC, same model files, same prompts. Upstream Strata 0.1.40.1 vs strata-rdna4, median of 3 runs: 📥 Prompt 32K: 515 → 1,826 tok/s 📥 Prompt 128K: 493 → 1,415 tok/s 📤 Output 32K: 51.5 → 84.7 tok/s ⏱️ First token at 1…
- Built on Qwen3.8 architecture with MoE (Mixture of Experts) and Flash-Next efficiency. Abliterated to remove refusal directions. GGUF format…
@HuggingModels · vorgestern · X · models, comparison
Built on Qwen3.8 architecture with MoE (Mixture of Experts) and Flash-Next efficiency. Abliterated to remove refusal directions. GGUF format for easy local inference. Vision-language capable. Optimized for llama.cpp. A p…
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
@coldniko · vorgestern · X · benchmark, hardware, release
Community benchmark attention on Ryzen AI Max / Strix Halo builds with Strata v0.1.40.
- Atomic Chat shipped a TensorRT integration into their engine last week. Qwen3.5-4B and Qwen3-8B through CUDA on a Blackwell RTX 5090: 3.43 s…
@0xPascual · vorgestern · X · models, release, gpu
Atomic Chat shipped a TensorRT integration into their engine last week. Qwen3.5-4B and Qwen3-8B through CUDA on a Blackwell RTX 5090: 3.43 seconds, against 10.49 seconds on llama.cpp. A 206% throughput gap on identical s…
- @steeve definitely not mistral. deepseek v4.1 flash, the fastest non-cerebras model in the world beats mistral large 4. qwen3.8 27b, SOTA fo…
@notnullptr · vorgestern · X · models, comparison
@steeve definitely not mistral. deepseek v4.1 flash, the fastest non-cerebras model in the world beats mistral large 4. qwen3.8 27b, SOTA for local people can actually run, almost matches it. qwen3.8-flash-next, runnable…
- @ReptileRaised @0xSero Qwen3.8-Flash-Next (125B/6B active + 51B N-gram) edges the dense 27B on most coding/agent benches (DeepSWE 58.7 vs 42…
@grok · vorgestern · X · benchmark, models, comparison
@ReptileRaised @0xSero Qwen3.8-Flash-Next (125B/6B active + 51B N-gram) edges the dense 27B on most coding/agent benches (DeepSWE 58.7 vs 42.2, JobBench 55.7 vs 33.4). On a 3090 with solid 4-bit/EXL3 quant + optimized st…
- Strata wins by a tight margin against all other Strix Halo inference engines in it’s first try release. Did not expect that! 🥳 https://t.co…
@coldniko · vorgestern · X · hardware, release, strata
- @coldniko @high52weeks
@1bitlabs · vorgestern · X · hardware, strata, comparison
@coldniko @high52weeks I had Grok @bot test Strata vs Halogen it's not apples to apples because it looks like Halogen maybe uses a custom quant to get the fast prefill? But the winner for me is still Halogen purely based…
- @Biggest I’m waiting @Alibaba_Qwen @QwenDevs 👀
@coldniko · vorgestern · X · models, comparison, strata
- @dvnold @MistralAI you can literally beat this with a tiny 129b Qwen3.8 Flash Next running on 1 local GPU, at much faster speed no less.
@_thomasip · vorgestern · X · benchmark, models, comparison
- ran the numbers on the hardware behind @primus_labs FHE engine vs Zama's, and the gap surprised me.
@Amir_rz_k · vorgestern · X · benchmark, gpu, pricing
ran the numbers on the hardware behind @primus_labs FHE engine vs Zama's, and the gap surprised me. Primus's AutoHoG hits 210+ confidential transfers/sec on 8x RTX 5090, consumer GPUs. At current street price (~$4,500/ca…
- ran the numbers on the hardware behind @primus_labs FHE engine vs Zama's, and the gap surprised me.
@Amir_rz_k · vorgestern · X · benchmark, gpu, pricing
ran the numbers on the hardware behind @primus_labs FHE engine vs Zama's, and the gap surprised me. Primus's AutoHoG hits 210+ confidential transfers/sec on 8x RTX 5090, consumer GPUs. At current street price (~$4,500/ca…
- For those who slept on Qwen3.8-Flash-Next — community ts-bench comparisons vs 27B stacks
@coldniko · vorgestern · X · models, benchmark, deploy
Highlights MoE local speed vs dense llama.cpp routes; links to models and install pages.
- @lxfater @redefinemee @ivanalog_com It's likely to be a model problem, not Strata problem. I run into a loop with ik_llama.cpp + 3.8 flash …
@wu89_j · vorgestern · X · models, strata, comparison
- @ggerganov Flash-Next support is the one I was waiting for. On my scavenged 2016 R730, single V100, llama.cpp gave me ~20 tok/s vs ~50 on St…
@VectorCrossProd · vorgestern · X · benchmark, models, strata
@ggerganov Flash-Next support is the one I was waiting for. On my scavenged 2016 R730, single V100, llama.cpp gave me ~20 tok/s vs ~50 on Strata. Gonna rerun on v0.6.0 and see if the gap closes. Less than pretty rack, bu…
- The real question with local Qwen3.8 inference isn’t:
@Oluwaphilemon1 · vorgestern · X · benchmark, models, strata
The real question with local Qwen3.8 inference isn’t: “Which model is smarter?” It’s: “Do I want the speed of Flash-Next, or the consistency of the 27B?” A new ts-bench comparison of Strata + Qwen3.8-Flash-Next vs llama.…
- every number in the map, sourced:
@yume_arasaki · vor 3 Tagen · X · benchmark, models, comparison
every number in the map, sourced: The total list would be insane, so i'm sure i'm missing some sources here : ( First the crazy one, it runs on 8GB of Vram: https://t.co/q924nO0h1M BENCHMARKS - Artificial Analysis, Flash…
- The same 111GB Qwen3.8-Flash-Next model file.
@Oluwaphilemon1 · vor 3 Tagen · X · benchmark, models, strata
The same 111GB Qwen3.8-Flash-Next model file. The same 2016 Xeon. The same single V100. But roughly twice the decode speed. Qwen3.8-Flash-Next UD-Q4_K_XL is getting around: ~20 tok/s with Strata versus ~9 tok/s with llam…
- @OpenAI @AnthropicAI Jackson__ measured vision: Strata median error 154 px vs llama.cpp 46 px – “the difference … is as large as the jump fr…
@ycnewsdigest · vor 3 Tagen · X · models, strata, comparison
- 🚨 NVIDIA LEAKS: RTX 5070 Ti vs. RX 9070 XT WAR!
@PrimeMediaSite · vor 3 Tagen · X · gpu, pricing, comparison
🚨 NVIDIA LEAKS: RTX 5070 Ti vs. RX 9070 XT WAR! NVIDIA’s latest RTX 5070 Ti vs. AMD’s RX 9070 XT. Who wins the GPU showdown? 🔥 NVIDIA leaks reveal massive price drops. Click... 📖 Read Full Story 👇 https://t.co/AqrZfQ…
Upgrades & migration
Version bumps, switch stories, and after-install notes from X.
- @IDFkushoverlord for dense models like qwen27B it was very slow for me even with all my efforts i could get preprocessing fast enough for my…
@l3d3ka · vorgestern · X · benchmark, models, strata
@IDFkushoverlord for dense models like qwen27B it was very slow for me even with all my efforts i could get preprocessing fast enough for my liking but since literally today i switched to strata inference engine and i’ve…
- Strata wins by a tight margin against other Strix Halo inference engines on first try release
@coldniko · vorgestern · X · benchmark, hardware, release
Community benchmark attention on Ryzen AI Max / Strix Halo builds with Strata v0.1.40.
- Strata wins by a tight margin against all other Strix Halo inference engines in it’s first try release. Did not expect that! 🥳 https://t.co…
@coldniko · vorgestern · X · hardware, release, strata
- @lxfater @redefinemee @ivanalog_com It's likely to be a model problem, not Strata problem. I run into a loop with ik_llama.cpp + 3.8 flash …
@wu89_j · vorgestern · X · models, strata, comparison
- Strix Halo officially supported in Strata v0.1.40
@coldniko · vor 3 Tagen · X · release, hardware, strata
AMD GPU improvements, multi-GPU batching, decode and Turing/Arc fixes — see release notes on site.
- tldr: if you're on 3090 + 64GB ram system, you should be running IQ3_S instead of IQ3_XXS.
@ItsmeAjayKV · vor 3 Tagen · X · models, strata, upgrade
tldr: if you're on 3090 + 64GB ram system, you should be running IQ3_S instead of IQ3_XXS. I switched to Qwen3.8-flash-next-GSQ-RSO-IQ3_S from IQ3_XXS since strata is fast. I noticed even tho model is bigger, the speed d…
- @ItsmeAjayKV @_ryu15_ so glad a follow you lol i have a 3090ti and 128gb of ddr4 and was running qwen 3.8 flash 2.5bpw on exl3..i just swi…
@BeeferShitlord · vor 3 Tagen · X · models, strata
@ItsmeAjayKV @_ryu15_ so glad a follow you lol i have a 3090ti and 128gb of ddr4 and was running qwen 3.8 flash 2.5bpw on exl3..i just switched to @support_huihui model on strata... jesus man i went from a ceiling of abo…
- https://t.co/ukYppIOjY8
@GameGPU_com · vor 3 Tagen · X · hardware, release, gpu
https://t.co/ukYppIOjY8 A PC user shared an upgrade from an Asus ROG Strix GeForce RTX 3090 to an XFX Mercury AMD Radeon RX 9070 XT Magnetic Air graphics card. Funds from selling the older flagship covered the new order …
- @coldniko
@simworld23 · vor 4 Tagen · X · benchmark, models, strata
@coldniko Strata v0.1.38 benchmark: Qwen3.8 Flash Next @ 131K ctx IQ1_M: 52.4 prompt / 37.9 output tok/s IQ2_XS: 70.2 prompt / 52.2 output tok/s Ryzen 9 7900X • 64GB DDR5 • RTX 3060 12GB IQ2_XS is looking very solid on a…
- @bnjmn_marie Switched to the Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS a few weeks ago, and have not looked back. It's totally comparable with deep…
@PJSmith · vor 5 Tagen · X · models, strata, comparison
@bnjmn_marie Switched to the Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS a few weeks ago, and have not looked back. It's totally comparable with deepseek flash 4.1 (in my work). Just run it on 'High'. 'Strata' sealed the deal. No…
- Tested Strata v0.1.38 with Qwen3.8-Flash IQ1_M on 2× RTX 3060s and 16GB RAM. At 64K context, Q4_0 KV read a 55K-token prompt at 486 tok/s vs…
@ghulamurtaza_ · vor 5 Tagen · X · benchmark, models, strata
Tested Strata v0.1.38 with Qwen3.8-Flash IQ1_M on 2× RTX 3060s and 16GB RAM. At 64K context, Q4_0 KV read a 55K-token prompt at 486 tok/s vs 459 on INT8 (+6%). Both answered correctly. INT8 decoded faster: 16.3 vs 14.9 t…
- @shazzadut @PaulineHansonOz You’re missing the point. It’s not about coffee. A lot of people count on the card to fill the gaps. Real estate…
@deboochatterjee · vor 6 Tagen · X · strata, pricing
@shazzadut @PaulineHansonOz You’re missing the point. It’s not about coffee. A lot of people count on the card to fill the gaps. Real estate offices have switched off card payments, so tenants can’t pay rent or bond by c…
- @viewhero1 Strata is definately a good move in the right direction. Llama.cpp is very unoptimized. I switched to a ninfer v100 fork. Llama I…
@DaayTerkErJerbs · vor 6 Tagen · X · benchmark, models, strata
@viewhero1 Strata is definately a good move in the right direction. Llama.cpp is very unoptimized. I switched to a ninfer v100 fork. Llama I got 2x 30 toks v100 32gb 27b. Ninfer I get 25 toks 8x. Strata is crushing llama…
- @Oluwaphilemon1 I just switched from Qwen 27b at 8bit. Now strata IQ3 and it’s wild. 90 tok/sec on dual 3090’s with NVlink and 128gb ddr4 …
@1secXrp · vor 6 Tagen · X · benchmark, models, strata
@Oluwaphilemon1 I just switched from Qwen 27b at 8bit. Now strata IQ3 and it’s wild. 90 tok/sec on dual 3090’s with NVlink and 128gb ddr4 ram. It even runs cooler as it spikes power between the cards and rarely sustains.…
- @joerg_peetz Thanks, dont laugh thats what I have 😄. I switched to bigger Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S with newest Strata engine, "IQ3_…
@Schweinert · vor 8 Tagen · X · models, strata, upgrade
@joerg_peetz Thanks, dont laugh thats what I have 😄. I switched to bigger Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S with newest Strata engine, "IQ3_XXS yields 80+ now". Found another flag that maybe can boost the speed even furt…
GPU quotes
| Model | Price | Date | Source |
|---|---|---|---|
| RTX 5090 | $1.999 (medium) | 2026-10-07 | @investies |
| RX 9070 XT | $700 (medium) | 2026-10-06 | @bottleneck_pc |
| RX 9060 XT 8GB | $299 (medium) | 2026-10-06 | @CoastalBendTech |
| RTX 5090 | $36 (medium) | 2026-10-06 | @0xoLand |
| RTX 5090 | $4.500 (medium) | 2026-10-06 | @Amir_rz_k |
| RTX 5090 | $4.500 (medium) | 2026-10-06 | @Amir_rz_k |
| RTX 5090 | $7,55 (medium) | 2026-10-06 | @Temperatur2com |
| RTX 3090 24GB | $3.450 (medium) | 2026-10-06 | @grok |
| RTX 5090 | $3 (medium) | 2026-10-06 | @chainbrain |
| RTX 5090 | $9.600 (medium) | 2026-10-05 | @DigitalNav |
| RTX 5080 16GB | $936,75 (medium) | 2026-10-05 | @GameGPU_com |
| RX 9070 XT | $5 (medium) | 2026-10-05 | @GameGPU_com |
| RTX 3090 | $899 (medium) | 2026-10-05 | @GameGPU_com |
| RTX 5090 | $7,71 (medium) | 2026-10-05 | @Temperatur2com |
| RTX 5090 | $1.999 (medium) | 2026-10-05 | @tobinaut |
| RTX 6000 | $16 (medium) | 2026-10-05 | @SebAaltonen |
| RX 9070 XT | $842,14 (medium) | 2026-10-04 | @Deals_Lantern |
| RX 9070 XT | $1.000 (medium) | 2026-10-04 | @MizoChris |
| RX 9070 XT | $2.813 (medium) | 2026-10-03 | @GameGPU_com |
| RTX 5070 | $850 (medium) | 2026-10-02 | @bottleneck_pc |
All signals (incl. GitHub fallback)
- @leypascua @HuggingModels X reports: 150-180 tok/s on RTX 5090 (Strata), 35-45 on 20-series for the Coder GGUF; 44 tps claimed on 32GB+ syst…
@grok · vorgestern · X · benchmark, models, strata
@leypascua @HuggingModels X reports: 150-180 tok/s on RTX 5090 (Strata), 35-45 on 20-series for the Coder GGUF; 44 tps claimed on 32GB+ systems. Reddit/HF: 29.6 GB resident works with SSD n-gram mmap. Coding benches hold…
- Old Ethereum mining GPUs got a new job: running a 125B model.
@ElmoVT · vorgestern · X · benchmark, models, strata
Old Ethereum mining GPUs got a new job: running a 125B model. Qwen3.8-Flash-Next + Strata, 16/24 GB GPUs, 64 GB RAM. Up to 138.7 tok/s. Real benchmarks, first-token latency, and cached turns. https://t.co/gFM9KnQeVq #Loc…
- 125B params at 94 tok/s on a 12GB RTX 5070. Strata runs Qwen 3.8 Flash Next on a gaming PC with 64GB RAM, Q2 quant. a year ago this needed a…
@arnabkarmkr · vorgestern · X · benchmark, models, strata
- @Androthon @Oluwaphilemon1 Qwen3.8-Flash-Next (125B MoE, ~6B active + 51B n-gram). Inference server is Strata. Needs 12GB+ VRAM and 64GB+ sy…
@grok · vorgestern · X · benchmark, models, strata
@Androthon @Oluwaphilemon1 Qwen3.8-Flash-Next (125B MoE, ~6B active + 51B n-gram). Inference server is Strata. Needs 12GB+ VRAM and 64GB+ system RAM for solid performance (32GB RAM minimum for lighter quants like Coder).…
- @Androthon @Oluwaphilemon1 It's just Strata. Strata is the complete open-source inference engine for Qwen3.8-Flash-Next. It handles expert c…
@grok · vorgestern · X · models, strata
@Androthon @Oluwaphilemon1 It's just Strata. Strata is the complete open-source inference engine for Qwen3.8-Flash-Next. It handles expert caching (VRAM/RAM/SSD), memory tiering, and speculative decoding on its own, then…
- Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090
@ItsmeAjayKV · vorgestern · X · models, strata, gpu
Time for test results of Qwen3.8-flash-next-GSQ-RSO-IQ3_S on my 3090 Here is a jungle fps theme game one-shot generated by IQ3_S System: - RTX 3090 - 64GB RAM - Strata Inference engine IQ3_S does produce better graphics …
- Single DGX Spark owners have a new model worth paying attention to.
@Oluwaphilemon1 · vorgestern · X · models, deploy
Single DGX Spark owners have a new model worth paying attention to. Qwen3.8-Flash-Next NVFP4 can now be served on one Spark with a setup that is surprisingly capable for a single 128GB unified-memory machine. The headlin…
- MLX-Serve 26.9.5 is starting to look less like a local model runner and more like an actual AI serving stack for Apple Silicon.
@Oluwaphilemon1 · vorgestern · X · models
MLX-Serve 26.9.5 is starting to look less like a local model runner and more like an actual AI serving stack for Apple Silicon. The headline additions are Qwen-Image-2.1 and Prism Bonsai 2. But the concurrency work is ar…
- @ItsmeAjayKV @0xSero @coldniko came here looking for this those tok/s numbers look anemic compared to what I'm experiencing with Strata
@rskjr · vorgestern · X · benchmark, strata, comparison
- NVIDIA launched its high-end RTX 50 "Blackwell" graphics cards, with the RTX 5090 as the flagship model.
@investies · vorgestern · X · release, gpu, pricing
NVIDIA launched its high-end RTX 50 "Blackwell" graphics cards, with the RTX 5090 as the flagship model. The RTX 5090 boasts impressive specs and a price tag of $1,999. What's the first game you'd play with this powerhou…
- Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers near…
@thetechnotice · vorgestern · X · benchmark, gpu, pricing
Forget the top-tier 5090. The PNY RTX 5080 Slim is on an incredible deal, offering massive value. Even better, the RTX 5070 Ti delivers nearly the same performance for a fraction of the price. This is the card to get! #G…
- Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strat…
@AgentWorkflowLa · vorgestern · X · benchmark, models, strata
Strata vs llama.cpp, same model and four GPUs: a 64K input took ~26s vs ~134s to first reasoning token. Decode was ~2-2.6x faster with Strata (MTP on). Cache reuse narrows the gap. Full benchmark: https://t.co/cpx6NOcU9i
- @LundukeJournal But it's a sad state if all the tokens are cloud based. Do at least some of the work on local LLMs. Strata is the best curre…
@Robert_of_Maine · vorgestern · X · strata
@LundukeJournal But it's a sad state if all the tokens are cloud based. Do at least some of the work on local LLMs. Strata is the best current inference engine for this if you've got at least 32GB of system ram.
- Prime Big Deal Days is live until 11:59pm PT tomorrow. New drops at midnight, 8am and 1pm Pacific.
@bottleneck_pc · vorgestern · X · gpu, pricing
Prime Big Deal Days is live until 11:59pm PT tomorrow. New drops at midnight, 8am and 1pm Pacific. The three lines to refresh for: RX 9070 XT under $700, Ryzen 7 9800X3D under $420, Corsair RM850e under $85. RAM and SSD …
- another few ns here
@CCDavidCode · vorgestern · X · benchmark, models, strata
another few ns here saved memory 8 times there. hmm dont need to wait on that Q8, plain decode llama.cpp: 22.0 tok/s Q8, MTP Early Q8 Strata prototype: 87.5 tok/s coding at 8K Strata now: 156.2 tok/s at 32K
- @_imdawon Strata is going to get faster
@CCDavidCode · vorgestern · X · benchmark, strata
- @coldniko Strix Halo 385 w/ 64gb memory getting 42 tp/s with strata IQ3_ XXS - on my llama.cpp build I was getting peak 27. Great job
@Prod_Philosophy · vorgestern · X · hardware, models, strata
- @IDFkushoverlord for dense models like qwen27B it was very slow for me even with all my efforts i could get preprocessing fast enough for my…
@l3d3ka · vorgestern · X · benchmark, models, strata
@IDFkushoverlord for dense models like qwen27B it was very slow for me even with all my efforts i could get preprocessing fast enough for my liking but since literally today i switched to strata inference engine and i’ve…
- 450+ tok/s of prefill on AMD Strix Halo.
@msala9 · vorgestern · X · hardware, benchmark, models
450+ tok/s of prefill on AMD Strix Halo. DS4 Halo accelerates DeepSeek V4 Flash on a single Radeon 8060S integrated GPU, in Strix Halo 128 GB. 454.59 tokens/s measured on a complete 4K prompt. Model: DeepSeek V4 Flash 07…
- @rskjr Hey Robert! I have seen your live stream and concerns in Strata and totally agree. Things are progressing a bit fast, just wanted to …
@coldniko · vorgestern · X · update, strata
- @JJismusic @MiaAI_lab Yes, that's an effective approach. EmbeddingGemma 2's 740M modular size loads easily alongside Qwen3.8-Flash-Next as a…
@grok · vorgestern · X · models
@JJismusic @MiaAI_lab Yes, that's an effective approach. EmbeddingGemma 2's 740M modular size loads easily alongside Qwen3.8-Flash-Next as a retrieval tool. Use it for unified search across text/code/images/video/audio, …
- @0xSero 5080 (16GB)
@hugovntr · vorgestern · X · benchmark, strata
- Strata is pushing Qwen3.8-Flash-Next to ~165 tok/s on my RTX 5090. 🤯
@XniX · vorgestern · X · benchmark, models, strata
Strata is pushing Qwen3.8-Flash-Next to ~165 tok/s on my RTX 5090. 🤯 524K context + Vision enabled, running a 125B-class MoE locally on a single consumer GPU. Seriously impressive. 🔥 #LocalLLM #Qwen #RTX5090
- Same PC, same model files, same prompts. Upstream Strata 0.1.40.1 vs strata-rdna4, median of 3 runs:
@vectorweft · vorgestern · X · benchmark, strata, comparison
Same PC, same model files, same prompts. Upstream Strata 0.1.40.1 vs strata-rdna4, median of 3 runs: 📥 Prompt 32K: 515 → 1,826 tok/s 📥 Prompt 128K: 493 → 1,415 tok/s 📤 Output 32K: 51.5 → 84.7 tok/s ⏱️ First token at 1…
- @coldniko's Strata runs a 125B MoE model on a desktop PC. I spent the last week tuning it for AMD RDNA4.
@vectorweft · vorgestern · X · benchmark, models, strata
@coldniko's Strata runs a 125B MoE model on a desktop PC. I spent the last week tuning it for AMD RDNA4. Qwen3.8-Flash-Next (Unsloth UD-Q4_K_XL) on 2x Radeon AI PRO R9700: ⚡ 1,826 tok/s prompt reading at 32K ⚡ 85-100 tok…
- @yume_arasaki Strata is an amazing project! But the benchmarks for the full precision model is NOT what you are getting with Strata. In my t…
@realjohnmock · vorgestern · X · benchmark, models, strata
@yume_arasaki Strata is an amazing project! But the benchmarks for the full precision model is NOT what you are getting with Strata. In my testing it's on par or slightly below with Qwen 3.8 27b at a similar tok/s rate.
- Built on Qwen3.8 architecture with MoE (Mixture of Experts) and Flash-Next efficiency. Abliterated to remove refusal directions. GGUF format…
@HuggingModels · vorgestern · X · models, comparison
Built on Qwen3.8 architecture with MoE (Mixture of Experts) and Flash-Next efficiency. Abliterated to remove refusal directions. GGUF format for easy local inference. Vision-language capable. Optimized for llama.cpp. A p…
- 125B MoE on a laptop: 55 → 105 tok/s in 3 days 🚀
@baltaci212 · vorgestern · X · benchmark, models, strata
125B MoE on a laptop: 55 → 105 tok/s in 3 days 🚀 RTX 5070 Ti Laptop (12 GB) + 64 GB RAM, Qwen3.8-Flash-Next Swift 1.5 IQ3_XXS on the Strata engine • 30-question set: 55 → 105 tok/s (30/30 correct) • Turkish answers: 94 …
- GLM 5.3 Flash on a single Tesla V100 32GB + 64GB of RAM
@PeasantSmith · vorgestern · X · benchmark, models, strata
GLM 5.3 Flash on a single Tesla V100 32GB + 64GB of RAM -18B Active parameters -25 tok/s peak decode -@UnslothAI IQ1_S (93GB file) -CPU I5-12600T +DDR5 5200 + Gen4 NVME -Custom strata The Volta GPU is probably the floor,…
- @ItsmeAjayKV They are preparing for enhanced n-grams most likely. Strata shows that it possible to run them faster and more efficient as wel…
@Bitcopath · vorgestern · X · benchmark, strata
@ItsmeAjayKV They are preparing for enhanced n-grams most likely. Strata shows that it possible to run them faster and more efficient as well. I'm telling you for some time, we'll able to use much larger models with our …
Methods
- X-first: users + comparison + upgrade search queries
- GitHub community (fallback, INTAKE_COMMUNITY_LIMIT)
- X API v2 (X_BEARER_TOKEN)
- RSSHub public mirrors (twitter/user, twitter/search)
- Self-host RSSHub recommended for rate limits — see docs/INTAKE.md
- DIYgod/RSSHub (twitter routes)
- github.com/vladkens/twscrape (session-based X scrape)
- github.com/zedeus/nitter (legacy RSS; instances unstable)