Issues / #610

#610 RTX A3000 / 0.1.38: a native IQ3_XXS decode window decomposes via STRATA_VERIFY_PROFILE, and its GPU side is memory-bandwidth-bound

closed · @yannickloth · 1 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quants

Description

## Context

Follow-up to #494 (the CPU-pool decode profile on this host). That issue sized the *host* side and left the GPU
side open: on the native IQ pack `--gpu-stages`, `--gpu-only-full` and `--no-pool`+`--spec` all fail, so the
window could not be decomposed. Engine **0.1.38**, host RTX A3000 12 GB (GA104, sm_86, 32 SMs, 3 MB L2, 192-bit
GDDR6 at 7 GHz = 336 GB/s), i7-12850HX, Swift 1.5 IQ3_XXS native pack (48 layers, 36 GDN + 12 QSA), serve stopped,
fixed `bench/e2e.sh` prompt.

## 1. The native window *can* be decomposed - through the serve path

`--gpu-stages` still refuses on a native pack (`session_replay_stages: not captured with the split`, because
per-layer graphs are skipped for native packs), but the serve path already prints the same information:

```
STRATA_VERIFY_PROFILE=1 STRATA_DECODE_TIMING=1  <engine> --serve <native args>
```

prints, per window, `strata decode GPU stages (ms/window): GDN layers: ... | QSA layers: ... | total ...` plus the
`verify (GPU-reach wait + per-layer host + stage)` split. This needed no code change. Suggest either documenting
this on the `--gpu-stages` path or making `--gpu-stages` fall back to the `Verifier` stamps for native packs - the
`prof_on_` timer is already there.

Fresh (mean of 3 runs) / warm (8th request of one session):

| term | fresh | warm |
| --- | ---: | ---: |
| GPU-reach wait | 26.3-27.1 ms | 24.1-27.6 ms |
| CPU expert pool (per-layer host) | 27.4-28.9 | 22.7-34.3 |
| tokens/window | 2.10 | 2.10-2.21 |
| reported cache hit rate | 52.6% | 67-70% |

Warm does **not** fall below fresh even though the pool does: at the higher hit rate more experts are computed on
the GPU, so the GPU term is roughly flat (the Amdahl correction to #494's "GPU 34% fresh / 62% warm").

## 2. The GDN projections are at the DRAM roofline

Ranked GPU compute (wait* excluded), ms/window fresh -> warm: `VRAM hits` (grouped int8 dp4a expert GEMV) 5.3 -> 6.9;
`hc-read1+router` 5.1 -> 5.2; GDN `q8+qkv gemv` 3.0 -> 3.2; `out-proj` 2.5 -> 2.7; `head` 2.5 -> 2.3; GDN `z`
2.0 -> 1.7; `shared+quant` 1.8 -> 2.0; QSA `q+q-idx` 1.7 -> 1.5, `attention` proper only ~0.4-0.6.

The GDN projection GEMVs (`attn_qkv`, `attn_gate`, `ssm_out`, all Q8_0 in this pack, 36 layers) read
**2.21 GB of weights per window** and measure 6.4 ms/window = **~345 GB/s** against the card's **336 GB/s** peak -
saturated. The native window is one captured CUDA graph, so the old "~13 tiny latency-bound launches per GDN
layer" no longer applies; what is left is execution at the bandwidth wall.

Things tried, both negative:

- The upstream multi-column MMVQ layout (`g_multi_exact = false`) is **not** a free win here: it changes the fixed
  prompt's tokens (it is equal to ncols=1 only to float rounding) **and** leaves every GDN stage unchanged
  (`q8+qkv` 3.15 -> 3.20, `z` 1.69 -> 1.73, `out-proj` 1.74 -> 1.83 ms/window). The exact default looks right for
  this GEMV class.
- `STRATA_GR_DOWN_MAX4=1` (0.1.38, bit-exact, +2.1% on RTX PRO 6000) is token-identical but **-0.43% e2e** on
  this host (5 x 6 warm requests/arm, interleaved), so it is a no-op here.

## Ask

1. Document (or wire up) the native stage profiler so others do not hit the same wall as #494.
2. For the GDN path, the only remaining lever we can see is **lower-precision projections** (Q8_0 -> Q5_K/IQ) - is
   that a direction you would take for the pack, or is GDN numerically sensitive enough to rule it out? Happy to
   measure a pack built that way.

Raw stage lines and the A/B output can be attached.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.