Issues / #1651

#1651 Add `swift` family support for `IQ3_S` (Swift 1.5 now has an IQ3_S tier)

open · @lplusc · 0 Kommentare · Auf GitHub

Setup & installNVIDIA / CUDAModels & quants

Beschreibung

## Summary

`IQ3_S` is currently restricted to the `qwen` family in `setup.py`, based on the
assumption that *"Swift 1.5 has no IQ3_S"*. That assumption is now out of date:
UkisAI released an **IQ3_S tier for Swift 1.5** on 2026-10-08. Please allow
`--family swift --model IQ3_S`.

## Evidence

`setup.py`:

```python
# the original model only (Swift 1.5 has no IQ3_S): matches the full BF16 model on the published benchmarks
"IQ3_S": {"about": "3.5-bit i-quant, the best quality (matches the full model), the slowest; needs a 64 GB PC "
                   "with little else running", "download_gb": 83.6, "ram_gb": 62, "arena_gb": 50.3,
          "families": ("qwen",)},
```

Because `families` defaults to `("qwen", "swift")`, the explicit `("qwen",)` on
`IQ3_S` is what blocks the Swift family. With the current release:

```text
--family swift  ->  --model choices: IQ2_XS, IQ3_XXS
--family qwen   ->  --model choices: Q2_0, IQ2_XS, IQ3_XXS, IQ3_S
```

## The files already exist and match your naming convention

Model: **https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF**

The two shards match the Swift pattern in `FAMILIES["swift"]`
(`Swift-Qwen3.8-Flash-Next-GSQ-RCO-{q}-0000{i}-of-00002.gguf`) exactly:

| File | Bytes | SHA-256 |
| --- | ---: | --- |
| `Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf` | 44,900,816,736 | `46cb3996cd1a17fe8bc5ae1f9bb16869480de5a03f17f868d6047a20a46edd3b` |
| `Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf` | 38,842,920,384 | `70ce6fb48a94fe34bd7e9a53a526be68066c64b5a5077cdc82fb31350b66b608` |

Combined size: **83.74 GB** (decimal). The BF16 `mmproj` is the same file already
shared by the other Swift tiers
(`mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf`,
SHA-256 `cd1140f4abba943ce8d9fa92fed8407e389954d26907efde7b72153c75bd8a60`).

Upstream reports IQ3_S as the lowest-KLD of the four tiers (development KLD
**0.144674**, vs 0.240139 for IQ3_XXS), so it is the quality-ceiling option for
this family.

## Request

1. Add `"swift"` to `IQ3_S`'s `families` (and drop the stale comment), so
   `--family swift --model IQ3_S` is selectable.
2. Please sanity-check the Swift-specific packaging constraints that already
   affected this family — e.g. the Swift `Q2_0` exclusion in the same dict
   (`#171`, one layer's experts split across the two shards) and the
   `--ple-gguf` / PLE-table placement, since these differ between tiers.
3. If IQ3_S's Swift shards turn out to need extra pack-tool handling, please
   keep it gated rather than silently enabling it.

## Environment (where this was hit)

- Strata `v0.1.41`
- 2× Tesla T10 16 GB (sm_75, Turing), 247 GiB RAM
- Weights already downloaded and SHA-256 verified against the official
  `SHA256SUMS` / `release-manifest.json`.

Happy to test a patch on this hardware and report back.

Mehr auf der Site

Links zu Install, Modellen, Releases.