Pull requests / #1402
#1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
open · @aflin · 0 comentarios · En GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descripción
## Title Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards Issue: none. Related: #824 (perronemirko's Qwen3.6 engine, closed by the history rewrite). That PR adds a separate `strata-qwen36` engine beside the main one; this one teaches the main engine the model instead, so the server, setup, the batched prompt path, MTP and the expert cache all apply to it, and the engine picks the model from the GGUF. ## Summary Qwen3.6-35B-A3B is the same family as Flash-Next (DeltaNet layers, full attention every fourth layer, a MoE with a shared expert) at 35B parameters: its experts take 12-14 GB of RAM instead of 23-50 GB. It is the model for PCs with 16-32 GB of RAM and 6-12 GB cards. Ornith-1.5-35B-A3B (ornith-ai, MIT) is a coding-agent fine-tune of it with the same architecture. The engine reads `general.architecture` and runs `qwen35moe` beside `qwen4exp` on one code path. Nothing on the command line says which model it is. Flash-Next's arithmetic is unchanged: its tokens are identical to upstream's own build (Coder IQ1_M, three prompts and an 8K prompt, the same expert cache size, before and after merging main). Measured on Linux (Debian 12), RTX 4070 Ti 12 GB, Ryzen 9 5950X, 62 GB RAM, CUDA 13.0; 32K context, 8-bit K/V, MTP on. "8 GB" is the same card with 4 GiB held by another process (emulated, not a real 8 GB card). | | Short chat, writes | 8K prompt, reads | 8K prompt, writes | | --- | ---: | ---: | ---: | | Qwen3.6 UD-IQ4_XS, 12 GB | 105 tok/s | 4,460 tok/s | 86 tok/s | | Qwen3.6 UD-IQ4_XS, 8 GB | 75 tok/s | 2,669 tok/s | 68 tok/s | | llama.cpp b11438, same file, 12 GB (`--fit on`) | 51 tok/s | 708 tok/s | 49 tok/s | | llama.cpp b11438, 8 GB | 39 tok/s | 490 tok/s | 36 tok/s | Greedy tokens match llama.cpp b11438 exactly on UD-IQ4_XS (three short prompts and the 8K prompt, 32 tokens each, expert cache at 3,000 slots). With a different cache size a near-tie can flip, because experts computed on the GPU round slightly differently from the CPU's (seen once: llama.cpp's top two at 28% and 26%). More numbers (UD-IQ3_S, Ornith, small cards, 256K) are in `docs/QWEN36.md`. ## What changed **Engine: model geometry** (`include/strata/core/layout.hpp`, `src/core/layout.cpp`, `gguf_reader.hpp`) - `ModelGeometry` gains the architecture, experts per token, and whether the model has PLE, the sparse-attention indexer and hyper-connections; `geometry_from_gguf` fills it and checks that each size fits the kernels' limits (all of Qwen3.6's are at or below Flash-Next's, which stay as compile-time capacities). **Engine: decode** (`src/core/verify.cpp`, `src/kernels/cuda/*.cu`, `src/kernels/cpu/*`, `src/core/expert_source.cpp`) - One residual stream (an RMSNorm before the mixer and before the MoE, the output added back), the `silu(z)` DeltaNet output gate, and dense attention: every cached cell is attended (`qsa_select_all`), with no indexer projections, snapshots or replay. - Kernel template instances for 8 or 12 query heads per K/V head, 32 or 48 DeltaNet heads, and 8 or 10 experts; the CPU pool and expert layout take the model's widths. - Native Q2_K, Q3_K and Q6_K expert formats (Unsloth's mixes use them); `iq_row_bytes` covers Q2_K (without it the MTP layer's Q2_K experts read as garbage). **Engine: prompt path** (`src/prefill/prefill.cpp`, `kernels.cu`) - The batched path for one stream: DeltaNet kernels for 32 heads and the silu gate, routing and combine for 8 of 256, a dense mode for the tensor-core prompt attention. NVIDIA sm_75+; elsewhere (`Prefill::supports`) the prompt is read through the decode windows. - After the merge: main's deferred shared expert uses the model's sizes; the opt-in CPU share (`STRATA_PREFILL_CPU_SHARE`) is off for one-stream models, since it is unmeasured there. **Engine: MTP** (`src/core/mtp.cpp`) - The draft layer is read from the model GGUF's `blk.40` (`eh_proj` split into its embedding and hidden halves, native projections and experts). No separate download. **Engine: VRAM on small cards** (`src/program/generate.cpp`, `Verifier::dense_attention_bytes`) - The cache sizing keeps the dense attention's buffers out of the expert cache. They grow with the context and are allocated after the cache: without this a 256K context failed on every card. - The small-card reserve floor is 500 MiB for models without the indexer (measured: 400 failed in the graph instantiation at 8K, 32K and 64K, 450 started) and stays 300 for Flash-Next. - A cache too small to work stops the start with how much is short and what makes room, instead of failing later on the prompt buffers. - Result: UD-IQ3_S with setup's settings starts from 4.17 GB free; with an 8K context and `--draft-vocab en`, from 3.9 GB. **Data** - `data/expert-profile-qwen36.bin`: 40 x 256 experts ranked by routing counts from 24 varied prompts (server `--expert-profile-save` with a 64-slot cache). Against an unranked profile, on prompts not among the 24 (three runs): 12 GB short-chat writes 99 -> 106 tok/s, 8K writes 84 -> 93 tok/s; no difference at 8 GB. **Setup** (`setup.py`) - Families `qwen36` (unsloth/Qwen3.6-35B-A3B-MTP-GGUF: UD-IQ4_XS, UD-IQ3_S) and `ornith` (bartowski/Ornith-1.5-35B-A3B-GGUF: IQ4_XS, IQ3_XXS), pinned revisions and SHA-256. ModelScope serves the same four files with the same SHA-256. - Suggested where no Flash-Next size fits the RAM (16-24 GB PCs). On a card under 5.5 GB setup recommends the smallest size and an 8K context (a recommendation, never forced). - The published engines (0.1.40 to 0.1.40.2) predate qwen35moe and their version cannot tell them apart, so setup looks for the support in the engine file itself and compiles the engine when a ready-made one lacks it. Flash-Next keeps the ready-made engine. - Tests: `tools/test_setup_qwen36.py`, `tools/test_setup_ornith.py`; updated `test_setup_choices.py`, `test_setup_unsloth.py`. **Tools** - `tools/iq_pack.py` leaves the MTP layer out of the expert table; `tools/make_profile.py --n-layer`; `tools/iq_fixture.py` adds Q2_K and Q4_0; `tools/strata_mcp.py` `recommend()` follows setup's rules. **Docs** - New `docs/QWEN36.md` (setup, what fits, small cards, speed, against llama.cpp, Ornith, differences, limits); sections in `docs/MODELS.md` and `README.md`. ## Extra Notes - **Tested on:** Linux (Debian 12, kernel 6.1), RTX 4070 Ti 12 GB (driver 580), CUDA 13.0, Ryzen 9 5950X, 62 GB RAM. **Not built or run on AMD (HIP) or Windows.** The prompt path for these models is NVIDIA-only (elsewhere prompts are read through the decode windows), and setup's compile path for them is untested on Windows. - **One GPU** only for these models (the layer split is not done for them); no images yet. - **Ornith's draft layer** is accepted less often than Qwen3.6's (61% vs 80% of all drafts on a chat prompt). A Q8_0 copy of the layer did no better, so it is the fine-tune's draft layer, not its quantization. - **Version:** QWEN36_ENGINE is 0.1.40.2 (this tree's version). CMakeLists and MIN_ENGINE are untouched, left to the release. - **Tests:** setup, MCP and profile tests pass, as do `iq_parity` (with and without `STRATA_PARITY_IQ_MMVQ`) and `native_grouped_parity`. `tools/test_setup_engine_hash.py` fails two tests; they fail the same way on main. - **Commits:** the work, and a merge of main at e8ca9af (0.1.40.2) with its 12 conflicts resolved. Squash or rebase as you prefer. Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.