Pull requests / #704

#704 STRATA_HC_Q8: hyper-connection projections as int8 + fp32 scale per 32

closed · @merbanan · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDAModels & quantsWindows

Description

Opt-in (STRATA_HC_Q8=1). The four BF16 hc_* projections per layer (1,200 MiB) are skipped by the loader and uploaded as int8 codes with one fp32 scale per 32 values (675 MiB); the freed VRAM goes to the expert cache.

The plain down/up kernels (single-token and multi-token, both tiles) read either form in the same uint4 register budget. With Q8 weights the multi read takes the split variant's down kernel (the plain one) instead of staged, the opt-in v3 read is skipped, and the unfused gr_read refuses. The prompt path expands the codes to a bf16 scratch for its GEMM. Not with a layer split.

RTX 2060 SUPER 8 GB, Q2_0, 128K context, k8v4 KV:
- expert cache 1,865 -> 2,262 slots (+397); decode 33.5-34.4 -> 34.7-36.0 tok/s
- teacher-forced, 3,070 tokens (code, prose, chat) through verify windows: top-1 94.2%, KL 3.5e-2, perplexity -0.050 +- 0.008 nats/token. Changing only the cache size (no HC_Q8) moves it as much: top-1 94.1%, KL 3.4e-2, +0.016 +- 0.008 - the difference is expert placement (GPU vs CPU rounding). gr_parity passes (the bf16 path is unchanged).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.