Issues / #513
#513 Feature Request: experimental support for Qwen3.8-Whittle-MoE-27B-A17.8B
closed · @Vitaliy86 · 1 Kommentare · Auf GitHub
Beschreibung
Would it be possible to consider adding experimental support for Qwen3.8-Whittle-MoE-27B-A17.8B to Strata? Model: https://huggingface.co/scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants I think this model could be particularly interesting for Strata because it is a relatively small MoE model compared with the larger Qwen3.8 models, while still having a substantial active parameter count. Why this model may be interesting Qwen3.8-Whittle-MoE-27B-A17.8B is a modified MoE architecture with: - 27B total parameters - approximately 17.8B active parameters - 64 routed experts - 16 experts selected per token - 1 shared expert - 64 layers - GGUF quantizations from Q2/Q3 up to Q8 The important part is that the model is already supported by recent llama.cpp as the "qwen35moe" architecture and does not require a custom llama.cpp fork or patches. For example, it can be loaded directly using llama.cpp: llama-server -hf scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants:Q4_K_S Could Strata support it? I understand that Strata is not simply a GUI wrapper around llama.cpp and that some of the Strata optimizations are specifically designed around the currently supported model architectures. Therefore, I am not suggesting that this would necessarily be as simple as adding another Hugging Face repository. The interesting question is whether Strata's existing: - expert loading / expert caching; - CPU/RAM ↔ GPU management; - expert offloading; - memory optimization; - model configuration; - speculative decoding / MTP infrastructure; could be adapted to the "qwen35moe" architecture used by Whittle. Even experimental support without MTP or other advanced features would already be very useful. Possible implementation approach Perhaps the first step could be to add it as a separate experimental model family, for example: Qwen3.8-Whittle-MoE-27B-A17.8B with automatic detection of the available GGUF quantization. The initial implementation could simply provide: 1. Hugging Face model download; 2. GGUF quant selection; 3. llama.cpp-based inference; 4. Strata's existing GPU/CPU memory management where applicable; 5. the existing Strata UI/API. Then the more advanced Strata-specific optimizations could be evaluated separately. MTP / speculative decoding I am particularly interested in whether Strata's existing MTP/speculative decoding approach could work with this model. I realize this may not be directly compatible because Whittle changes the MoE structure compared with the original Qwen3.8 model. So MTP is not a requirement for the initial integration. It would be interesting simply to know whether it is technically possible. Why I think this could be useful This model has much lower total memory requirements than the large Qwen3.8 MoE models. For example, the available GGUF variants are roughly: - Q3_K_M — ~14 GB - Q4_K_S — ~16 GB - Q4_K_M — ~17 GB - Q5_K_M — ~20 GB - Q6_K — ~23 GB - Q8_0 — ~29 GB This makes the model potentially interesting for systems with 16–24 GB GPUs, especially if Strata's expert offloading/caching techniques can reduce VRAM requirements and improve performance. In short Could you please investigate whether Qwen3.8-Whittle-MoE-27B-A17.8B could be supported by Strata, even initially as an experimental model? The model already works with recent llama.cpp using the "qwen35moe" architecture, so perhaps the main challenge is adapting Strata's model-specific expert management and optimizations rather than implementing a completely new inference backend. Even if full support with MTP and all Strata optimizations is not currently practical, I would be very interested in seeing whether basic support could be added first and the advanced features evaluated afterwards. Model: https://huggingface.co/scima/Qwen3.8-Whittle-MoE-27B-A17.8B-GGUF-all-quants Thank you for considering it.
Mehr auf der Site
Links zu Install, Modellen, Releases.