Pull requests / #1465
#1465 cuda: retain HC norm products in existing scratch
open · @W1nge · 0 comentarios · En GitHub
Multi-GPUAMD / HIPNVIDIA / CUDAModels & quantsWindows
Descripción
The split/staged HC normalization reads the residual twice and recomputes a pending write after its reduction. Store the unscaled R * norm product in the existing xn scratch during the first pass, then scale it in place. The reduction order and two multiplication roundings stay unchanged. This extracts the useful part of our local stream-norm path without adding another kernel variant or tuning flag. Validation on RTX 2080 Ti (sm_75), Windows/MSVC, driver 560.94: - The existing fused_gr_selftest passed all plain/split/staged comparisons, including windows of 1..8 tokens with and without a pending write, bit for bit. - A direct norm comparison also matched xn and rs bit for bit at T=1,2,4,8, both write modes. - A CUDA-graph microprobe (1000 launches per graph, 10 alternating rounds) measured reductions from 3.2% to 17.1% for this norm kernel. These are component timings, not end-to-end decode gains. Other GPU generations and HIP are not measured; draft pending broader validation. The existing per-device correctness check/fallback remains intact. The existing internal selftest can be run in a small CUDA TU including src/kernels/cuda/fused_gr.cu, calling fused_gr_selftest(ok, why), and requiring ok[1], ok[2], ok[3]. This checks the actual production kernels. No changes to public APIs, allocation sizes or hardware selection. Based directly on d5ea713, independent of the prefill memory and prompt-reuse PRs. Validation update (2026-10-08): a complete current-main integration build including #1451, #1454, #1455, #1457 and #1465 passed four real IQ3_XXS model baseline/integration pairs on RTX 2080 Ti: default, GR_UNFUSED, PREFILL_BF16X2 and legacy RING_BYTES modes. Each used INT8 KV, 194 input tokens, two prefill chunks and 8 output tokens; all runs exited successfully and output IDs matched within each pair. A 256K configured-context startup and short generation also passed (this did not fill a 256K prompt). These are combined regression checks, not an isolated performance result or exhaustive state equality. Other hardware/backends remain untested. The earlier draft-only status reflected the absence of any current-main model run; now requesting review with these limits explicit.
En el sitio
Enlaces a install, modelos, releases.