Issues / #556
#556 EXL3 (exllamav3 trellis) weight support
closed · @teee9 · 1 comments · View on GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Description
Feature request: let the engine run EXL3 (exllamav3's trellis quantization) checkpoints directly. For this model's size class the (supported) GGUF options top out at IQ3_S, and on my workload the 3-bit quality drop isn't acceptable. exllamav3's 4.05 bpw EXL3 is the best size/quality tradeoff I've found that fits my RTX 3090 24 GB + 64 GB DDR5 Windows 11 machine.
**Why**
- I currently serve `Qwen3.8-Flash-Next-exl3-4.05bpw_H6_NG6` (with BF16 ngram + vision tower, FP16 KV, 65536 context) on TabbyAPI/exllamav3. I have experimentally let Deepseek v4.1 Flash write a local exllamav3 patch that records and seeds CPU/GPU expert placement (CPU expert split with static hot-first placement), seeded from Strata's shipped `data/expert-profile.bin` expert ranking, plus MTP with 3 draft tokens. That took stock exllamav3 CPU offload from ~20-22 tok/s to ~26-30 tok/s production, ~360 tok/s prompt.
- It doesn't go further on this hardware. Profiling showed decode at this size is per-layer CPU-GPU handoff-latency-bound, not RAM-bandwidth-bound: moving more experts to the GPU didn't raise decode, and cutting the CPU expert tail's compute further didn't either. The remaining gap is the engine loop, which Strata has and exllamav3 doesn't.
- I've also tried the 3.05 bpw EXL3 weights and 3-bit GGUF; both were just not up to my standards. The ask is a format, not a smaller quant.
**What the format is (from exllamav3 source, MIT)**
- Per-tensor safetensors: `.trellis` (int16 bit windows; 16×16 tile = `16*K[+8]` u16, K inferred from tensor width), `.suh`/`.svh` (fp16 per-channel scales), optional `.mul1`/`.mcg` marker selects the codebook. No header; no fused `gate_up_proj`.
- This checkpoint: 48 trunk MoE layers plus the MTP layer, 512 experts each, × {gate,up,down} = 75,264 expert trellis tensors, **all `mul1`**; 323 trunk trellis tensors; BF16 token embedding; PLE/ngram is one `[320001536 × 61]` int16 trellis (~36 GB, separate file); MTP head separate.
- Decode = bit-window extraction + procedural mul1 codebook + **mandatory Hadamard-128 input/output** (weights live in a rotated basis) with `suh`/`svh` folded in.
- All decode kernels are plain CUDA/C++ (no CUTLASS); `moe_mul1.cpp` implements exactly this codebook with AVX2/AVX-512 tiers; the GPU core is `codebook.cuh` + `exl3_dq.cuh`.
Would native EXL3 weight support be something the project would consider adding?Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.