Issues / #505
#505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
closed · @lawsirlawsir-png · 5 comentários · No GitHub
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
## Platform (all values verified on the machine)
| Item | Value |
|---|---|
| GPU | AMD Radeon PRO **W7800 48 GB** (gfx1100 / RDNA3 / wave32) |
| CPU | Ryzen 5 7500F (6 cores, AVX-512 + VNNI) |
| RAM | **30 GiB** DDR5-6000 dual-channel |
| OS / stack | Linux, ROCm **7.2.4**, hipBLASLt **1.2.2 (100202)**, PCIe Gen4 x16 |
| Engine | local HIP build, gfx1100 offload-arch; production binary sha256 `b3bcb0f5…` (0.1.34 build, 2026-10-02); earlier campaign on 0.1.27/0.1.29 |
| Model | `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS` (2 shards, HF revision `ed59f92…`) |
**Engine args (production, unchanged since 09-30):**
```
--pack packs/iq3_xxs --native …-00001.gguf --ple-gguf …-00002.gguf
--mmap-experts --expert-profile data/expert-profile.bin --expert-cache auto
--prefill auto --spec 3 --spec-min-p 0.7 --mtp mtp/rt
--max-context 262144 --kv k8v4 --pool-workers 5 --pcie-frac 0
--vram-reserve-mib 2048 --adapt-every 0 --vision
```
Env: `STRATA_PREFILL_MMQ=1`, `STRATA_PREFILL_RING=96`, `STRATA_IO_THREADS=32`, `STRATA_HIPBLASLT_TUNING=tools/hip/gfx1100-hipblaslt-100202.txt` (self-calibrated, 26 rows).
🔴 **All speed numbers below are measured with thinking OFF** (`chat_template_kwargs: {"enable_thinking": false}`). Per `docs/DETAILS.md`, the model default is **high** — so these are not directly comparable with the official rates tables unless the bench's thinking level is pinned on both sides. (One GSM8K arm ran with thinking on; it is listed separately as a throughput reference only.)
## Headline: 48 GB VRAM full-resident is a regime the official tables never measured
The official IQ3_XXS reference (RTX 5070 12 GB + 64 GB RAM) assumes experts spill to host/SSD. Here the expert cache auto-sizes to **24,576 slots / 39.97 GiB** with **100% decode hit rate** — experts never leave VRAM. In this regime:
### 1. Production workload, engine 0.1.34 (server log, mechanical extraction, n=573 arms)
Real interactive session, **no thinking**, context accumulated to **80–120K tokens**:
| Metric | Measured |
|---|---|
| **decode @ ctx 80–120K** | **mean 76.5 tok/s** (min 45.3 / max 100.7) |
| expert cache hit rate | 93.1–97.0% |
| draft acceptance | 0.76–0.94 (e.g. 320/336, 671/793) |
| fresh prefill (18–20K prompt, 0 reused, same log) | 1,196–1,231 tok/s |
| fresh prefill (160K prompt) | 1,043.8 tok/s |
**decode at 100K context (76.5) exceeds the official IQ3_XXS 4K reference (61.6) by ~24%.** We did not run the controlled `STRATA_HIP_ROUTER_OLD=1` A/B (interrupted), so this is observational — but it suggests the 0.1.31–0.1.34 AMD-line gains (`#262` intrinsics, 0.1.32 `hip_router_fast`) also benefit **gfx1100**, not only RDNA4.
Also: **no decode cliff from 4K→100K** in our config (`--kv k8v4`, `--spec 3`) — decode actually rises (55→76) because long arms have higher draft acceptance (0.85–0.94 vs 0.63–0.76 in bench). Another gfx1100 report (7900 XTX, 24 GB) saw a 32K cliff attributed to the default 32K MTP draft window. **Could you confirm the MTP draft-window default in 0.1.34 and whether `--kv k8v4` changes it?**
### 2. Official `bench_prefill.py`, engine 0.1.27 + self-calibrated table (conditions matched exactly: 4 fresh prompts 140/280/140/280 rules + 4 follow-ups, 128 max_tokens, temp 0, seed 42, `--prefill auto --kv int8 --mtp mtp/rt --expert-cache auto`)
| kind | prompt tokens | fresh | prefill t/s | decode t/s |
|---|---:|---:|---:|---:|
| fresh | 4,210 | 4,210 | **1,085.2** | 55.2 |
| fresh | 8,830 | 8,830 | **1,198.0** | 54.7 |
| fresh | 4,210 | 4,210 | **1,177.1** | 56.0 |
| fresh | 8,830 | 8,830 | **1,176.9** | 56.4 |
| follow-up avg | — | 248 | 343.8 | 51.9 |
vs official IQ3_XXS (RTX 5070): prefill 4K **108–117%**, decode 4K **90%**.
### 3. 256K context (`--kv k8v4`)
- 160,062-token prompt end-to-end: HTTP 200, 154.3 s, **prefill 1,043.8 tok/s, decode 57.9 tok/s**
- 128K (`--kv int8`) → 256K (`k8v4`): decode 55.0 → **56.6 (+2.9%)**, prefill 1,151.4 → **1,204.4 (+4.6%)**, resident experts −3.1%
- 300,059-token prompt correctly rejected with HTTP 400 ⇒ 262144 genuinely enforced
- k8v4 "pays off mostly on large cards at long contexts" — confirmed; on a 48 GB card it is nearly free (KV fully in VRAM, `--kv-resident` not needed)
## Four findings worth documenting
1. **hipBLASLt table version gate silently costs 75% prefill.** Shipped tables are 100100/100200; ROCm 7.2.4 ships hipBLASLt 1.2.2 (**100202**). Mismatch prints one stderr line (`… version mismatch: file=100200 runtime=100202; using hipBLASEx`) and prefill drops **622 → 1,088 tok/s** when fixed. A community 7900 XTX user on ROCm 7.2.3 independently hit 100202; PR #170 added a 100202 table but it appears absent from `tools/hip/` in 0.1.33. **Suggestion:** ship a 100202 table for gfx1100, or make setup print an explicit "run tools/hip/tune_hipblaslt.cpp" hint on mismatch. Our tuner command (works verbatim):
```sh
/opt/rocm-7.2.4/llvm/bin/clang++ -x hip -O2 -std=c++20 --offload-arch=gfx1100 \
--rocm-path=/opt/rocm-7.2.4 -I/opt/rocm-7.2.4/include tools/hip/tune_hipblaslt.cpp \
-L/opt/rocm-7.2.4/lib -lhipblaslt -lhipblas -lamdhip64 \
-Wl,-rpath,/opt/rocm-7.2.4/lib -o /tmp/tune_hipblaslt
ARGS=$(awk 'NR>2 {printf "--case %s,%s,%s,%s,%s ", $1,$5,$2,$3,$4}' tools/hip/gfx1100-hipblaslt-100200.txt)
/tmp/tune_hipblaslt --workspace-mib 32 $ARGS --tuning-out tools/hip/gfx1100-hipblaslt-100202.txt
```
2. **Decode variance is driven by MTP acceptance rate, which the docs never mention.** Same server, same session, `decode expert cache hit rate: 100.0%` (77,760/83,520): acceptance 0.63→49.8 t/s, 0.67→51.0, 0.72→56.0, 0.76→56.4. Docs say "2.4–3.2 tokens per pass" without an acceptance-rate baseline; we measured 1.65–1.73 tokens/pass during bench. **Suggestion:** publish per-model draft acceptance alongside the rates table so users can tell cache misses from drafting variance.
3. **IQ2_XS is 30% SLOWER than IQ3_XXS on this machine — opposite of the official table.** Strict interleaved A/B (same host/script/args, 8 runs each, fixed prompt wording): decode IQ3_XXS **54.0 ± 2.5** vs IQ2_XS **41.5 ± 0.6** (IQ3 faster 27–36% across 1K/4K/16K); prefill differs only +1.7%. We believe the official IQ2_XS row (78.6 @4K) comes from **Swift 1.5** (shorter-thinking fine-tune, different output-length distribution) — if so, the table should label the model variant, because users comparing original Flash-Next get the opposite ranking. We ruled out 13 candidates (config diff, RAM speed, CPU ISA, cache hit rate, draft quality, `STRATA_NO_IQ4NL`, `STRATA_CPU_YMM`, PR #241, PR #242 — gap even widened to 35.1%). With experts 100% GPU-resident, CPU-side knobs (`STRATA_IQ_MT_MIN`) can't explain it. **Open question for you: what in the single-token decode path makes IQ2_XS slower on gfx1100?**
4. **Vision on gfx1100 via llama.cpp `GGML_HIP=ON` — one CMake flag, not a port.** Upstream 0.1.33 still says "no HIP encoder build yet". We built `tools/vision/strata_vision.cpp` with `cmake -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1100 -DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++` (llama.cpp pin `3cf03257` already in `_deps`): READY 6.0 s vs CPU 10–30+ min; per image 4,000 tokens / 5.3 s vs CPU 480 tokens / 12 s (~30×); peak VRAM +1.61 GiB (inside `--vram-reserve-mib 2048`); end-to-end image QA correct through the OpenAI endpoint. Worth formalizing as an official build option.
## Parameter-level confirmations (for the FAQ)
| Finding | Data |
|---|---|
| `--expert-cache auto` must not be hand-set | manual 2048: 5.08 GiB cache, prefill 2.8 t/s; `auto`: 42.52 GiB, 6.9 t/s (8.4×) |
| `--mtp mtp/rt` is not optional | without: decode 18.92, acceptance 0.005; with: 46.43, 0.731 (2.45×) |
| `--spec` design ceiling is 3 | spec 3: 56.2–56.7; spec 4: 53.8–54.6 |
| `--pool-workers` 2→5 no change | 361.4 vs 323.1 t/s; host stream ~546 ms constant ⇒ 6-core CPU is not the prefill bottleneck |
| shard 2 is shared by 4 quants | IQ2_XS/IQ3_S/IQ3_XXS/Q2_0 shard-2 SHA256 identical ⇒ hardlink saves 28.8 GB per quant |
| `serve.server` startup line is misleading | prints "loading the experts into RAM (about 47 GB)" unconditionally; on 30 GB RAM the `--mmap-experts` path is what actually runs |
## Throughput reference (thinking ON, conc 4, n=200 GSM8K, seed 42, greedy, max_tokens 3072)
Strata Flash-Next IQ3_XXS: 190/200 = 95.0%, mean 394.9 completion tokens (incl. `reasoning_content`), **50.68 t/s aggregate**. Listed only as a workload reference — not comparable with the no-thinking speed numbers above.
## Method notes
- Prefill comparisons use **fresh (0 reused)** arms only; cached arms discarded per `bench_prefill.py` rules.
- Model-vs-model A/B uses fixed prompt wording + interleaved runs (we previously saw 24% decode swing from prompt wording alone — that was our own methodology bug, fixed).
- Cross-card comparisons (W7800/ROCm/Linux vs RTX 5070/Windows) are positional, not controlled A/B.
- 0.1.31–0.1.34 AMD gains are **not claimed** for gfx1100 here; §1 is observational evidence only.
Happy to run any controlled follow-up on this box (e.g. `STRATA_HIP_ROUTER_OLD=1` vs default at 4K/16K, or re-running `bench_prefill.py` at a pinned `reasoning_effort` to align with the official tables).
---No site
Links install, modelos, releases.