Issues / #1651
#1651 Add `swift` family support for `IQ3_S` (Swift 1.5 now has an IQ3_S tier)
open · @lplusc · 0 comentarios · En GitHub
Setup & installNVIDIA / CUDAModels & quants
Descripción
## Summary
`IQ3_S` is currently restricted to the `qwen` family in `setup.py`, based on the
assumption that *"Swift 1.5 has no IQ3_S"*. That assumption is now out of date:
UkisAI released an **IQ3_S tier for Swift 1.5** on 2026-10-08. Please allow
`--family swift --model IQ3_S`.
## Evidence
`setup.py`:
```python
# the original model only (Swift 1.5 has no IQ3_S): matches the full BF16 model on the published benchmarks
"IQ3_S": {"about": "3.5-bit i-quant, the best quality (matches the full model), the slowest; needs a 64 GB PC "
"with little else running", "download_gb": 83.6, "ram_gb": 62, "arena_gb": 50.3,
"families": ("qwen",)},
```
Because `families` defaults to `("qwen", "swift")`, the explicit `("qwen",)` on
`IQ3_S` is what blocks the Swift family. With the current release:
```text
--family swift -> --model choices: IQ2_XS, IQ3_XXS
--family qwen -> --model choices: Q2_0, IQ2_XS, IQ3_XXS, IQ3_S
```
## The files already exist and match your naming convention
Model: **https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF**
The two shards match the Swift pattern in `FAMILIES["swift"]`
(`Swift-Qwen3.8-Flash-Next-GSQ-RCO-{q}-0000{i}-of-00002.gguf`) exactly:
| File | Bytes | SHA-256 |
| --- | ---: | --- |
| `Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf` | 44,900,816,736 | `46cb3996cd1a17fe8bc5ae1f9bb16869480de5a03f17f868d6047a20a46edd3b` |
| `Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf` | 38,842,920,384 | `70ce6fb48a94fe34bd7e9a53a526be68066c64b5a5077cdc82fb31350b66b608` |
Combined size: **83.74 GB** (decimal). The BF16 `mmproj` is the same file already
shared by the other Swift tiers
(`mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf`,
SHA-256 `cd1140f4abba943ce8d9fa92fed8407e389954d26907efde7b72153c75bd8a60`).
Upstream reports IQ3_S as the lowest-KLD of the four tiers (development KLD
**0.144674**, vs 0.240139 for IQ3_XXS), so it is the quality-ceiling option for
this family.
## Request
1. Add `"swift"` to `IQ3_S`'s `families` (and drop the stale comment), so
`--family swift --model IQ3_S` is selectable.
2. Please sanity-check the Swift-specific packaging constraints that already
affected this family — e.g. the Swift `Q2_0` exclusion in the same dict
(`#171`, one layer's experts split across the two shards) and the
`--ple-gguf` / PLE-table placement, since these differ between tiers.
3. If IQ3_S's Swift shards turn out to need extra pack-tool handling, please
keep it gated rather than silently enabling it.
## Environment (where this was hit)
- Strata `v0.1.41`
- 2× Tesla T10 16 GB (sm_75, Turing), 247 GiB RAM
- Weights already downloaded and SHA-256 verified against the official
`SHA256SUMS` / `release-manifest.json`.
Happy to test a patch on this hardware and report back.
En el sitio
Enlaces a install, modelos, releases.