Issues / #1387

#1387 [Feature Request] Support for higher quantizations (Q8 / unquantized) for Qwen 3.8 / 4 with high-end hardware (128GB VRAM / 192GB RAM)

open · @Bobsosa · 2 comentarios · En GitHub

Multi-GPUModels & quants

Descripción

### Feature Request

First of all, thank you for the incredible work on Strata! The performance you've unlocked for Qwen3.8-Flash-Next on consumer hardware is truly mind-blowing.

While Strata currently does an amazing job optimizing large models for 12GB/16GB VRAM configurations using lower quantizations (like Q4), I am wondering if there are plans to extend its capabilities for high-end or multi-GPU setups. 

I am currently running a system with **128GB of VRAM** and **192GB of system RAM**. With this hardware headroom, I would love to run larger versions or higher-quality quantizations (such as Q8 or even unquantized/BF16) of the **Qwen 3.8** family, and eventually the upcoming **Qwen 4** models, while still leveraging Strata's lightning-fast MoE expert caching mechanism.

### Proposed Enhancement
Are there any plans or architectural possibilities to:
1. **Support higher quantizations** (e.g., Q8, FP16/BF16) within Strata's framework?
2. **Optimize expert offloading/caching** specifically for setups with massive VRAM/RAM pools, allowing larger or multiple experts to remain in the cache simultaneously?
3. Ensure future compatibility with **Qwen 4** high-precision variants?

I would love to hear your thoughts on whether supporting high-end hardware profiles fits into the roadmap for Strata. Thank you again for this game-changing project!

En el sitio

Enlaces a install, modelos, releases.