Issues / #1463

#1463 MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)

open · @mathcuei · 0 comentários · No GitHub

BenchmarksSetup & installServer & APIAMD / HIPModels & quantsDocumentation

Descrição

Results of a test of Strata **0.1.40.1** (`82f46a8`) on one **AMD Instinct MI50 32 GB (gfx906)**, measured on 2026-10-06. This complements the existing MI50 result (`bench/results/2026-10-04-community-mi50`, 0.1.38, another host, `--spec 3`): a later engine version, an AVX-512 host, and it adds 252K, a run with the card limited to 16 GB, and temperature readings. Runs were driven by Claude Code on my own machine. I am posting it as an issue with the numbers and configuration inline (no files attached); the raw logs and scripts are available if you want them as a `bench/results` folder.

**Short version:** decode **37-55 tok/s** on short prompts, **44.4 / 41.0 tok/s at 34K / 68K** of context, prefill **409-428 tok/s** from 8K to 119K tokens. With the full 32 GB the three-needle check passed at **126K (3/3)** and **252K (3/3)**, decode **42.5 tok/s at 126K** and **36.3 tok/s at 252K**. With the card limited to 16 GB it passed at 126K (31.2 tok/s) but **failed at 252K: the answer was `!!!!...` (0/3)**. That one was not reproduced (see "The 252K / 16 GB failure"). One run per case, synthetic text: these numbers do not establish answer quality or performance on other workloads.

## Hardware and software

- MI50 32 GB (Vega 20, gfx906), 34,342,961,152 bytes of VRAM; the engine's PCIe probe: 25.1 GB/s host to device. No power cap set by me, clocks not changed.
- Intel Core i9-11900H (the OS reports it as "Genuine Intel(R) CPU 0000 @ 2.60GHz", up to 4.8 GHz), 8 cores / 16 threads, AVX-512; 62 GB RAM; Ubuntu 26.04.1 LTS, kernel 7.0.0-38-generic. 7 expert-pool workers on logical processors 1-7, host thread on 0.
- ROCm **7.2.2** (AMD clang 22.0.0git, `roc-7.2.2`); `lld` needed a compatible `libxml2` on `LD_LIBRARY_PATH` on this OS.
- Strata `v0.1.40.1`, commit `82f46a8c8f475f001ad76d92f58f4a4f8ffb0253`, built with `-DSTRATA_HIP_GFX906=ON` as in `docs/AMD_HIP.md`; the engine reports version 0.1.40.
- Nothing else used the GPU; no system setting (power, driver, kernel) was changed.

## Model and configuration

`ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, `IQ2_XS` (the GGUF already on the machine; pack and MTP layer built by the setup flow). Vision off, temperature 0, no reasoning. Engine arguments (paths removed), 32 GB / 131,072:

```
--pack <data-dir>/packs/iq2_xs --native <models-dir>/.../IQ2_XS/...-00001-of-00002.gguf --ple-gguf <models-dir>/.../...-00002-of-00002.gguf
--expert-profile <strata-dir>/data/expert-profile.bin --expert-cache auto --prefill auto
--spec 4 --spec-min-p 0.5 --mtp <data-dir>/mtp/rt --max-context 131072 --kv int8 --kv-resident 32768
```

The 262,144 runs use `--max-context 262144`; the "16 GB" runs add `--vram-reserve-mib 16384` (= 32,752 - 16,368, what a 16 GB MI50 exposes), so they are **a simulation by Strata's reserve, not a real 16 GB card**. `--kv-resident 32768` (KV streaming) is what the installer chose; I did not compare against it being off.

- **32 GB:** 19,473 of 24,576 experts (~79%, 26.17 GiB) in VRAM, pre-filled from the profile, no eviction; decode cache hit rate 95.8-98.5%; ~39 GB of RAM in use; loading the 33.02 GiB of experts took ~1 min (3.68 GiB/s).
- **16 GB:** 8,099 slots (~33%, 10.85 GiB); hit rate 73-91%.

## Results (single runs; the engine's own `strata serve:` timings)

| Case | 32 GB | 16 GB (simulated) |
|---|---:|---:|
| Decode, short prompt (4 requests) | 37.2 / 54.2 / 55.1 / 40.6 tok/s (MTP drafts accepted 47-76%) | 30-34 tok/s |
| Decode at 34K context (350 tokens) | 44.4 tok/s | 35.7 tok/s |
| Decode at 68K context (350 tokens) | 41.0 tok/s | 38.9 tok/s |
| Prefill, 26-55-token prompts | 63-79 tok/s | 43-44 tok/s |
| Prefill, 8K to 119K tokens | **409-428 tok/s** (client-side; 410-430 server-side) | **394-419 tok/s** |
| Needle in the middle, 2K to 119K tokens | 6 of 6 | 6 of 6 |

Three 6-digit needles at 10 / 50 / 90 % depth, prompt near each limit; the second request re-sends the same text with another question and reuses the first one's prefix (114,688 tokens at 126K, 245,760 at 252K), so the last column is decode, not prefill:

| Mode | Context | Prompt tokens | Prefill | Needles | Decode with the full context |
|---|---|---:|---:|---|---:|
| 32 GB | 131,072 | 125,991 | 409 tok/s (308 s) | **3 of 3** | **42.5 tok/s** (hit 98.3%) |
| 32 GB | 262,144 | 252,304 | 370 tok/s (681 s) | **3 of 3** | **36.3 tok/s** (hit 98.5%) |
| 16 GB | 131,072 | 125,991 | 403 tok/s (313 s) | **3 of 3** | **31.2 tok/s** (hit 91.2%) |
| 16 GB | 262,144 | 252,304 | 382 tok/s (660 s) | **0 of 3: `!!!!...`** | 256 tokens at 26.8 tok/s, no draft accepted |

<details><summary>Engine request lines for the four long runs</summary>

```
32 GB, 126K : prompt 125991 tokens = 0 reused + 125991 read in 307796 ms (409.3 tok/s), 23 generated in 524 ms (43.9 tok/s), drafts accepted 15 of 20
              prompt 125985 tokens = 114688 reused + 11297 read in 31579 ms (357.7 tok/s), 293 generated in 6894 ms (42.5 tok/s), drafts accepted 144 of 236
32 GB, 252K : prompt 252304 tokens = 0 reused + 252304 read in 681378 ms (370.3 tok/s), 23 generated in 6606 ms (3.5 tok/s), drafts accepted 16 of 18
              prompt 252298 tokens = 245760 reused + 6538 read in 21504 ms (304.0 tok/s), 300 generated in 8265 ms (36.3 tok/s), drafts accepted 141 of 255
16 GB, 126K : prompt 125991 tokens = 0 reused + 125991 read in 312521 ms (403.1 tok/s), 23 generated in 7644 ms (3.0 tok/s), drafts accepted 15 of 20
              prompt 125985 tokens = 114688 reused + 11297 read in 31823 ms (355.0 tok/s), 300 generated in 9618 ms (31.2 tok/s), drafts accepted 134 of 252
16 GB, 252K : prompt 252304 tokens = 0 reused + 252304 read in 660486 ms (382.0 tok/s), 60 generated in 3576 ms (16.8 tok/s), drafts accepted 3 of 3
              prompt 252298 tokens = 245760 reused + 6538 read in 20185 ms (323.9 tok/s), 256 generated in 9563 ms (26.8 tok/s), drafts accepted 0 of 0
```
</details>

## Temperatures

Junction from `rocm-smi` every 10 s; memory temperature was **not** recorded. Prefill of 8K-119K reached **99-101 C junction at 155-200 W** (edge 78 C) in the 32 GB mode and 89-97 C in the 16 GB mode; the 126K runs stayed at 89-97 C and the 252K runs reached 97-104 C (the 32 GB 252K run peaked at 103 C and passed). Back to 40-53 C at idle. My guard (stop at 104 C twice in a row) fired once, during a repeat of the 16 GB 252K run (the request ended with HTTP 503). The card's sysfs lists critical limits of 105 C (junction/edge) and 94 C (HBM).

## The 252K / 16 GB failure

Output `!!!!...` and 0 of 3 needles at 252,304 tokens; the server logged no error and `dmesg` showed no GPU reset. The same prompt passed in the full 32 GB mode. The repeat was cut by my temperature guard, so I do not know whether it reproduces. The same `!!!!` pattern at long context is reported on other cards (#606, #879, #871); I found no earlier gfx906 report. I have a **hypothesis, not verified**, that heat in the HBM was involved (junction reached 102-104 C in both 16 GB 252K runs); there is no direct evidence, and a different expert cache (8,099 slots) may matter too.

## Notes for other MI50 users

- On ROCm 7.2.2, v0.1.40.1 did **not build** with `-DSTRATA_HIP_GFX906=ON` until three small fixes: a bare `return;` in a `bool` function in `fused_gr.cu`, and two `#if` guards in `mtp.cpp` and `vmm.cpp`. I looked at `main` (d5ea71337): all three are already fixed there. I have not rebuilt or re-run `main`.
- `setup.py` has no `gfx906` in `AMD_ARCHS` (checked on `main`), so the installer does not recognise the card: I wrote `engine/BUILD.json` by hand (`source: local-hip-gfx906`, `archs: [gfx906]`, the source hash) so setup accepted the engine.
- The full card held 19,473 of 24,576 experts at both 131,072 and 262,144; the 252K prefill ran at 370 tok/s, about 10% below the 126K one.

## Limits

One quant (IQ2_XS), one run per case, synthetic repetitive text (simple needles), no `--parallel`, no soak, no power limit, no memory-temperature log. The 16 GB rows come from a reserve, not a real 16 GB card. The numbers come from the engine version and host above.

No site

Links install, modelos, releases.