Pull requests / #1262
#1262 STRATA_HC_REQ8: hyper-connection projections requantized to int8 + fp32 scale per 32 at load (rebased #704)
open · @merbanan · 0 commentaires · Sur GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Description
The rebased #704, under the name asked for there (`STRATA_HC_Q8` now reads the GGUF's Q8_0 bytes, a different thing). Opt-in, off by default. ## What The four BF16 `hc_*` projections per layer (1,200 MiB) are skipped by the loader and uploaded as int8 codes with one fp32 scale per 32 values (675 MiB). The freed VRAM goes to the expert cache like any other. The changes requested on #704: - **Q8 is a template parameter** of the down/up kernels (single- and multi-token: `gr_down_kernel`, `gr_up_kernel`, `gr_down_multi_kernel`, `gr_up_multi_kernel`). The BF16 instantiations compile exactly as before; the launch sites pick the Q8 ones when `w_down` / `w_up` carry scales. The inject rows stay BF16. - **64-bit seek** in the loader (`fseeko` / `_fseeki64`). - **A layer split warns and keeps BF16** instead of exiting. - **KL check**: below. Paths that read the BF16 rows directly are skipped for it: the staged and v3 down kernels, `STRATA_HC_UPMIX`, `STRATA_PF_HCDOWN`, `STRATA_HCD_EXACT`, and the gfx906 `STRATA_GR_SPLIT` path. The prompt path expands the codes into one reused bf16 scratch for its GEMM. ## Measured RTX 2060 SUPER 8 GB (sm_75), Ryzen 9 3900X, 60 GB RAM, Q2_0, `--max-context 8192 --kv int8`, 400 MiB reserve (the distribution check's engines below): | | expert cache slots | dense weights loaded | | --- | ---: | ---: | | BF16 (default) | 2,044 | 1,416 MiB | | `STRATA_HC_REQ8=1` | 2,441 (+397) | 216 MiB + 675 MiB int8 | ### Distribution check (teacher-forced) The method of `docs/UNSLOTH_Q4.md`: a serve engine with `STRATA_LOGPOS` + `STRATA_LOGPOS_TOPK=256`, `--short-read 1600` so every prompt token is read through the verify windows and scored, `--adapt-every 100000 --pcie-frac 0` (fixed experts), greedy. Three ~1,000-token texts of this repository (code: `native_rope.cu`; docs: `DETAILS.md`; prose: `BATCHING.md`). KL is bf16 || other over the bf16 run's top 256 plus one bucket for the rest. Control: a second BF16 run. | comparison | text | positions | argmax same | top-10 overlap | KL mean | KL median | KL p99 | | --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | | bf16 vs bf16 (a second run) | all three | 3,327 | 100% | 100% | 0 | 0 | 0 | | bf16 vs REQ8 | code | 1,084 | 93.3% | 91.1% | 0.069 | 0.0077 | 1.03 | | bf16 vs REQ8 | docs | 1,145 | 99.7% | 92.2% | 0.0063 | 0.00017 | 0.066 | | bf16 vs REQ8 | prose | 1,098 | 97.7% | 95.4% | 0.0066 | 0.0012 | 0.088 | | bf16 vs REQ8 | all | 3,327 | | | 0.027 | 0.0011 | 0.43 | The BF16 reruns are bit-identical, so the REQ8 rows are the int8 projections alone. For scale: the prompt-attention kernel's check (`bench/results/2026-10-03-v100-prompt-attn`) moved the median KL by 0.0007-0.022 with controls that only change the rounding order, and the experimental speed projection (`bench/results/2026-09-27-esp`) measured mean 0.063. Code is the most sensitive of the three texts here (its heavy tail: p99 1.03, max 6.9). Reproduce: `bench/results/2026-10-06-hc-req8/req8_kl.py run`, then `compare` (paths at the top of the script). 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sur le site
Liens install, modèles, releases.