Issues / #610
#610 RTX A3000 / 0.1.38: a native IQ3_XXS decode window decomposes via STRATA_VERIFY_PROFILE, and its GPU side is memory-bandwidth-bound
closed · @yannickloth · 1 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quants
Description
## Context Follow-up to #494 (the CPU-pool decode profile on this host). That issue sized the *host* side and left the GPU side open: on the native IQ pack `--gpu-stages`, `--gpu-only-full` and `--no-pool`+`--spec` all fail, so the window could not be decomposed. Engine **0.1.38**, host RTX A3000 12 GB (GA104, sm_86, 32 SMs, 3 MB L2, 192-bit GDDR6 at 7 GHz = 336 GB/s), i7-12850HX, Swift 1.5 IQ3_XXS native pack (48 layers, 36 GDN + 12 QSA), serve stopped, fixed `bench/e2e.sh` prompt. ## 1. The native window *can* be decomposed - through the serve path `--gpu-stages` still refuses on a native pack (`session_replay_stages: not captured with the split`, because per-layer graphs are skipped for native packs), but the serve path already prints the same information: ``` STRATA_VERIFY_PROFILE=1 STRATA_DECODE_TIMING=1 <engine> --serve <native args> ``` prints, per window, `strata decode GPU stages (ms/window): GDN layers: ... | QSA layers: ... | total ...` plus the `verify (GPU-reach wait + per-layer host + stage)` split. This needed no code change. Suggest either documenting this on the `--gpu-stages` path or making `--gpu-stages` fall back to the `Verifier` stamps for native packs - the `prof_on_` timer is already there. Fresh (mean of 3 runs) / warm (8th request of one session): | term | fresh | warm | | --- | ---: | ---: | | GPU-reach wait | 26.3-27.1 ms | 24.1-27.6 ms | | CPU expert pool (per-layer host) | 27.4-28.9 | 22.7-34.3 | | tokens/window | 2.10 | 2.10-2.21 | | reported cache hit rate | 52.6% | 67-70% | Warm does **not** fall below fresh even though the pool does: at the higher hit rate more experts are computed on the GPU, so the GPU term is roughly flat (the Amdahl correction to #494's "GPU 34% fresh / 62% warm"). ## 2. The GDN projections are at the DRAM roofline Ranked GPU compute (wait* excluded), ms/window fresh -> warm: `VRAM hits` (grouped int8 dp4a expert GEMV) 5.3 -> 6.9; `hc-read1+router` 5.1 -> 5.2; GDN `q8+qkv gemv` 3.0 -> 3.2; `out-proj` 2.5 -> 2.7; `head` 2.5 -> 2.3; GDN `z` 2.0 -> 1.7; `shared+quant` 1.8 -> 2.0; QSA `q+q-idx` 1.7 -> 1.5, `attention` proper only ~0.4-0.6. The GDN projection GEMVs (`attn_qkv`, `attn_gate`, `ssm_out`, all Q8_0 in this pack, 36 layers) read **2.21 GB of weights per window** and measure 6.4 ms/window = **~345 GB/s** against the card's **336 GB/s** peak - saturated. The native window is one captured CUDA graph, so the old "~13 tiny latency-bound launches per GDN layer" no longer applies; what is left is execution at the bandwidth wall. Things tried, both negative: - The upstream multi-column MMVQ layout (`g_multi_exact = false`) is **not** a free win here: it changes the fixed prompt's tokens (it is equal to ncols=1 only to float rounding) **and** leaves every GDN stage unchanged (`q8+qkv` 3.15 -> 3.20, `z` 1.69 -> 1.73, `out-proj` 1.74 -> 1.83 ms/window). The exact default looks right for this GEMV class. - `STRATA_GR_DOWN_MAX4=1` (0.1.38, bit-exact, +2.1% on RTX PRO 6000) is token-identical but **-0.43% e2e** on this host (5 x 6 warm requests/arm, interleaved), so it is a no-op here. ## Ask 1. Document (or wire up) the native stage profiler so others do not hit the same wall as #494. 2. For the GDN path, the only remaining lever we can see is **lower-precision projections** (Q8_0 -> Q5_K/IQ) - is that a direction you would take for the pack, or is GDN numerically sensitive enough to rule it out? Happy to measure a pack built that way. Raw stage lines and the A/B output can be attached.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.