Pull requests / #1457
#1457 prefill: reuse shared scratch for the HC read
open · draft · @W1nge · 0 Kommentare · Auf GitHub
NVIDIA / CUDAModels & quantsWindows
Beschreibung
The HC read's normalized rows, low-rank projections and gate currently retain separate allocations throughout prefill. Reuse the existing attention/MoE scratch region for the complete HC read. A shared take_hc helper supplies both the allocation layout and byte estimate, including the unfused FP32 copy, optional padding and BF16 low parts. HC reads finish before attention/MoE uses the region on the compute stream. The fused next-half normalization follows those consumers; the existing fusion guard disables it before PLE, which also borrows the region. Persistent residuals, row scales, injection and mixed outputs stay separate. The legacy ring-byte estimate remains conservative. At T=32768, with default BF16 HC mode and no padding, the removed separate payloads total 1980 MiB (640 MiB normalized rows, 60 MiB low-rank buffers, 1280 MiB gate). Actual savings include allocator rounding and depend on the largest scratch set. No new hardware-specific switch or chunk default. Validation on RTX 2080 Ti / Windows: - Both ordinary and native/MMQ/fused prefill translation units compile. - Complete engine builds on current main plus this PR and #1451/#1454/#1455/#1465. - Four real-model baseline/integrated pairs pass: default, STRATA_GR_UNFUSED=1, STRATA_PREFILL_BF16X2=1, STRATA_RING_BYTES=0. Each uses IQ3_XXS, INT8 KV, 194 input tokens, chunk 128 (two prompt chunks), and 8 output tokens. All exit successfully with matching output IDs per pair. These are short combined-regression checks, not isolated speed measurements or exhaustive state/logit equality. Long prompts, control vectors, padding, multi-device and non-CUDA runtime coverage remain outstanding; keep as draft. Based directly on d5ea713. Independent of the other PRs; integration with #1454 needs the adjacent allocation/count edits combined (already tested locally).
Mehr auf der Site
Links zu Install, Modellen, Releases.