Pull requests / #1660

#1660 prefill: let native-layout packs take STRATA_MMQ_RESIDENT_SORT_NE

open · @ruibeikaa · 0 Kommentare · Auf GitHub

Multi-GPUNVIDIA / CUDAModels & quants

Beschreibung

The row-count sort of a fully resident layer's experts (c6a0d403) skips native-layout packs, though slot/src/off are built the same way for both layouts and nothing reads off[e + 1] as a count. This drops that one condition; the switch stays opt-in.

Measured on 4× GP100 (2× PH402, sm_60), Swift 1.5 IQ3_XXS served native with Q8_0 dense shards, `--layer-split 12,24,36`, every expert resident. Switch off → on, the same build, replies identical bit for bit:

| prompt read | off | on |
|---|---|---|
| 1.2K fresh | 7.27 s | 5.32 s |
| 1.1K appended at 35K | 7.95 s | 5.87 s |
| 35K fresh | 63.3 s | 54.1 s |

The expert GEMMs take 44% less time; nothing else moves, and decode is unchanged. Only sm_60 was checked for identical output; whether newer architectures' MMQ (stream-k) keeps the bits under the reordered groups is not measured.

Mehr auf der Site

Links zu Install, Modellen, Releases.