Pull requests / #844

#844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)

closed · @architectds · 0 comentários · No GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindows

Descrição

Adds Unsloth's `UD-Q3_K_XL` (revision `38bb39e`, three shards, 90.0 GB) to setup as a third Unsloth size, handled
exactly like UD-IQ4_XS: `START-HERE.bat --setup --family unsloth --model UD-Q3_K_XL`.

**No engine work is needed.** Despite its name the file has no Q3_K experts: the experts' down rows are 640 values
long, which K-quants' 256-value blocks do not divide. From the shard headers:

| Tensors | Format |
| --- | --- |
| experts' gate/up | IQ3_XXS in 47 layers, IQ4_XS in layer 2 |
| experts' down | IQ4_NL in 43 layers, Q8_0 in layers 2, 4, 30, 46 and 47 |
| attention, shared experts, hyper-connection projections | Q8_0 (UD-IQ4_XS's dense side) |
| head / PLE table | Q6_K / IQ4_NL |

The engine has kernels for all of them; only setup refused the file by name.

**Changes**
- `setup.py`: the model entry (three pinned shards with sizes and SHA-256 from the Hub's LFS pointers, engine 0.1.38 or
  newer, the RAM budget, `--compat-bf16`, images asked as for UD-IQ4_XS), the family text, the supported-files text
  and the `--resident-budget-gib` help. The size menu becomes UD-IQ4_XS (still the default), UD-Q3_K_XL, UD-Q4_K_XL
  (experimental, last).
- Docs: a UD-Q3_K_XL section in `docs/UNSLOTH_Q4.md` with the formats, the hashes and the measurements below;
  `docs/MODELS.md`, `README.md`, `docs/AI_SETUP.md`.
- Tests: `test_setup_unsloth.py` (a `Q3KXL` class: pins, budget, install, `--check`, engine version; the menus with
  three sizes), `test_setup_choices.py` (the supported-files text, a UD-Q3_K_XL name).

**Measured** on one PC (2026-10-04): RTX 3060 12 GB + RTX 5070 Ti 16 GB (both PCIe 3.0 x8), Ryzen 9 5900XT, 64 GB
DDR4-2933, against the GSQ-RCO IQ3_S with the same engine and settings for both. The engine was a 0.1.39 fork
(architectds/Strata, with two-GPU changes) and the file ran in the low-RAM mapped mode on a layer split
(`iq_pack.py --compat-bf16 --experts-bin`, `--mmap-experts`):

| | GSQ-RCO IQ3_S | UD-Q3_K_XL |
| --- | ---: | ---: |
| Same next token as the official Qwen 3.8 Flash API (31 answers, 14,960 tokens, read teacher-forced) | 93.0% | 93.0% |
| 5-token KL against the API's top 5 | 0.0536 | 0.0519 |
| Six trap and code questions, thinking on, two seeds each | 12 of 12, 11,048 tokens | 12 of 12, 10,620 tokens |
| Writing, in an agent-like session (100K start, eight 1-5K turns) | 77.9 tokens/s | 56.6 tokens/s |
| Reading that session's 100K-token start | 61.6 s | 81.8 s |

As close to the official model as IQ3_S on these checks, and slower on this PC (its experts are 11% larger, so fewer
fit in VRAM).

**Not run:** setup's own path for this file (the stock engine, one GPU, a RAM budget, the GGUF read in place). It is
UD-IQ4_XS's path; the docs say so and ask for reports.

**Tests:** every `python tools/test_setup_<name>.py` passes (18 files); `test_iq_pack.py` 23 OK.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.