Issues / #1479

#1479 HIP: STRATA_HEAD_MIX_MULTI / STRATA_ONE_TOKEN_COMMIT are hard-off; as opt-ins head-mix is exact and +0.4% decode on R9700 (patch inside)

open · @jkuepker · 0 commentaires · Sur GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsLinux

Description

Three paths ported from eddoursul's fork (F7) are compiled to `return false` on HIP:

- `head_mix_multi_enabled()` in `src/core/verify.cpp` (`STRATA_HEAD_MIX_MULTI`)
- the `head_mix_multi_on` gate in `MtpDrafter::record_rest` in `src/core/mtp.cpp`
- `one_token_self_commit()` in `src/core/verify.cpp` (`STRATA_ONE_TOKEN_COMMIT`)

That means nobody can try them on AMD without a rebuild. I patched them into **HIP opt-ins** (`=1` turns them on; the default stays off on HIP and CUDA is unchanged) and measured them on an R9700 (gfx1201):

| | decode tok/s (A/B/A/B) | answers |
|---|---|---|
| default | 91.6 / 91.6 | — |
| `STRATA_HEAD_MIX_MULTI=1` | **92.0 / 91.9 (+0.4%)**, faster on each of the 3 prompts in both rounds | bitwise identical |
| `STRATA_ONE_TOKEN_COMMIT=1` | 91.6 (0) | bitwise identical |
| both | 92.0 | bitwise identical |

Decode was measured with 3 fixed greedy prompts, 300 tokens and 3 rounds through the server, with a spread of 0.1-0.2% between server starts.

The gain is small, but it is exact and free. A default-on for gfx12 (or at least the opt-in) seems reasonable. Patch:

```diff
diff --git a/src/core/mtp.cpp b/src/core/mtp.cpp
index e73ce24..27ac0ea 100644
--- a/src/core/mtp.cpp
+++ b/src/core/mtp.cpp
@@ -852,7 +852,8 @@ bool MtpDrafter::record_rest(int step_row, cudaStream_t cs, std::string& err) {
         }();
         static const bool head_mix_multi_on = [] {
 #if defined(STRATA_USE_HIP)
-            return false;
+            const char* v = std::getenv("STRATA_HEAD_MIX_MULTI");   // HIP: opt-in (=1)
+            return v != nullptr && std::atoi(v) != 0;
 #else
             const char* v = std::getenv("STRATA_HEAD_MIX_MULTI");
             return v == nullptr || std::atoi(v) != 0;
diff --git a/src/core/verify.cpp b/src/core/verify.cpp
index 6b80569..30ac512 100644
--- a/src/core/verify.cpp
+++ b/src/core/verify.cpp
@@ -126,7 +126,11 @@ bool sh_stream_on() {
 // STRATA_HEAD_MIX_MULTI=0 does too): bitwise the same sums.
 bool head_mix_multi_enabled() {
 #if defined(STRATA_USE_HIP)
-    return false;
+    static const bool on = [] {
+        const char* v = std::getenv("STRATA_HEAD_MIX_MULTI");   // HIP: opt-in (=1)
+        return v != nullptr && std::atoi(v) != 0;
+    }();
+    return on;
 #else
     static const bool on = [] {
         const char* v = std::getenv("STRATA_HEAD_MIX_MULTI");
@@ -138,7 +142,11 @@ bool head_mix_multi_enabled() {
 
 bool one_token_self_commit() {
 #if defined(STRATA_USE_HIP)
-    return false;
+    static const bool on = [] {
+        const char* v = std::getenv("STRATA_ONE_TOKEN_COMMIT");   // HIP: opt-in (=1)
+        return v != nullptr && std::atoi(v) != 0;
+    }();
+    return on;
 #else
     static const bool on = [] {
         const char* v = std::getenv("STRATA_ONE_TOKEN_COMMIT");
```

**Not included:** the same treatment for `STRATA_MTP_SHARED_BRANCH` in `mtp.cpp` does not work on HIP. With `=1`, every request fails with `mtp: the shared expert's branch`, because the `cudaEventRecord(sh_fork_)` / `cudaStreamWaitEvent(side_)` pair fails, and `sh_fork_` / `side_` look like they are never created on HIP. That gate has to stay `false` unless the side stream is ported.

**Environment:** Strata v0.1.40.3 (d5ea713), Linux, AMD Radeon AI PRO R9700 (gfx1201, 32 GB), Ryzen 9 9900X, 30 GB RAM, TheRock ROCm 7.14.0a20260612 (`STRATA_ROCM_VERSION`), Coder IQ1_M, `--resident-experts`, `STRATA_HIP_WMMA=1`, `--prefill auto:32768`, all 12,288 experts in VRAM.
**Quality check used for every variant:** greedy needle retrieval at ~4K / ~16K / ~57K tokens plus a code task. All 4 pass, and the 4 answer hashes are identical to the baseline.

Sur le site

Liens install, modèles, releases.