Pull requests / #108
#108 prefill: bit-identical kernel speed-ups (batched indexer append, parallel GDN conv, column-split GDN recurrence, one-launch embedding gather)
closed · @q8atnight · 0 comentários · No GitHub
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quants
Descrição
## What Four prompt-path kernels that were launch- or latency-bound, rewritten so that **every output stays bit-identical** (each change has an env switch back to the old path and a debug self-check that runs both from the same state and compares every byte): | change | before | after | switch / check | |---|---|---|---| | QSA indexer append | one 1-block launch per token (8,192 per QSA layer per chunk) | 2 launches per chunk (`native_qsa_indexer_append_chunk`: each block does what the per-token kernel does at a completing position; placeholders/tail written by a finishing kernel) | `STRATA_PF_IDX_BATCH=0` / `STRATA_PF_IDX_CHECK=1` | | GDN 4-tap conv | 80 blocks, one thread per channel walking all T tokens | one thread per (token, channel) + a tiny history kernel (same expression) | `STRATA_PF_CONV_PAR=0` / `STRATA_PF_CONV_CHECK=1` | | GDN recurrence | 48 blocks (one per value head, 16-warp barriers), loads on the critical path | software-pipelined loads, and split over value **columns** (192 blocks of 4 warps; each column's FMA chains / reductions unchanged; the output RMS norm is finished by a second kernel from the same per-warp partial sums) | `STRATA_GDN_PIPELINE=0`, `STRATA_PF_REC_SPLIT=0` / `STRATA_PF_REC_CHECK=1` | | token embeddings (GGUF-form table) | one dequant launch per token | one `gather_dev` launch per chunk (same dequantizer); picture rows uploaded over it | `STRATA_PF_EMBED_BATCH=0` | ## Measured RTX 3090, IQ3_S, 34K-token prompt (`STRATA_PREFILL_TIMING=1`, GPU timeline of the prompt path): - qsa indexer 1,082 ms → 2 ms, gdn conv 501 → 183 ms, gdn recurrence 1,671 → 1,188 ms, embedding host time 180 → 0 ms - ~2 s less per 34K prompt; in our setup the prompt rate went 1,680 → 1,851 tok/s from these four alone. The self-checks printed `IDENTICAL` for every layer and chunk, including chunk boundaries (p0 = 8192, 16384) and a continuation starting at a non-4-aligned position (p0 = 9229). Measured on 2x RTX 3090 / Ryzen 9 3950X / 121 GB RAM (the numbers above are the primary card's own work, so they apply to a single card too). Developed with an AI coding assistant; all numbers measured on the hardware above.
No site
Links install, modelos, releases.