Pull requests / #844
#844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)
closed · @architectds · 0 コメント · GitHub で見る
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows
本文
Adds Unsloth's `UD-Q3_K_XL` (revision `38bb39e`, three shards, 90.0 GB) to setup as a third Unsloth size, handled exactly like UD-IQ4_XS: `START-HERE.bat --setup --family unsloth --model UD-Q3_K_XL`. **No engine work is needed.** Despite its name the file has no Q3_K experts: the experts' down rows are 640 values long, which K-quants' 256-value blocks do not divide. From the shard headers: | Tensors | Format | | --- | --- | | experts' gate/up | IQ3_XXS in 47 layers, IQ4_XS in layer 2 | | experts' down | IQ4_NL in 43 layers, Q8_0 in layers 2, 4, 30, 46 and 47 | | attention, shared experts, hyper-connection projections | Q8_0 (UD-IQ4_XS's dense side) | | head / PLE table | Q6_K / IQ4_NL | The engine has kernels for all of them; only setup refused the file by name. **Changes** - `setup.py`: the model entry (three pinned shards with sizes and SHA-256 from the Hub's LFS pointers, engine 0.1.38 or newer, the RAM budget, `--compat-bf16`, images asked as for UD-IQ4_XS), the family text, the supported-files text and the `--resident-budget-gib` help. The size menu becomes UD-IQ4_XS (still the default), UD-Q3_K_XL, UD-Q4_K_XL (experimental, last). - Docs: a UD-Q3_K_XL section in `docs/UNSLOTH_Q4.md` with the formats, the hashes and the measurements below; `docs/MODELS.md`, `README.md`, `docs/AI_SETUP.md`. - Tests: `test_setup_unsloth.py` (a `Q3KXL` class: pins, budget, install, `--check`, engine version; the menus with three sizes), `test_setup_choices.py` (the supported-files text, a UD-Q3_K_XL name). **Measured** on one PC (2026-10-04): RTX 3060 12 GB + RTX 5070 Ti 16 GB (both PCIe 3.0 x8), Ryzen 9 5900XT, 64 GB DDR4-2933, against the GSQ-RCO IQ3_S with the same engine and settings for both. The engine was a 0.1.39 fork (architectds/Strata, with two-GPU changes) and the file ran in the low-RAM mapped mode on a layer split (`iq_pack.py --compat-bf16 --experts-bin`, `--mmap-experts`): | | GSQ-RCO IQ3_S | UD-Q3_K_XL | | --- | ---: | ---: | | Same next token as the official Qwen 3.8 Flash API (31 answers, 14,960 tokens, read teacher-forced) | 93.0% | 93.0% | | 5-token KL against the API's top 5 | 0.0536 | 0.0519 | | Six trap and code questions, thinking on, two seeds each | 12 of 12, 11,048 tokens | 12 of 12, 10,620 tokens | | Writing, in an agent-like session (100K start, eight 1-5K turns) | 77.9 tokens/s | 56.6 tokens/s | | Reading that session's 100K-token start | 61.6 s | 81.8 s | As close to the official model as IQ3_S on these checks, and slower on this PC (its experts are 11% larger, so fewer fit in VRAM). **Not run:** setup's own path for this file (the stock engine, one GPU, a RAM budget, the GGUF read in place). It is UD-IQ4_XS's path; the docs say so and ask for reports. **Tests:** every `python tools/test_setup_<name>.py` passes (18 files); `test_iq_pack.py` 23 OK. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
関連リンク
インストール・モデル・リリースへの站内リンク。