Issues / #1479
#1479 HIP: STRATA_HEAD_MIX_MULTI / STRATA_ONE_TOKEN_COMMIT are hard-off; as opt-ins head-mix is exact and +0.4% decode on R9700 (patch inside)
open · @jkuepker · 0 评论 · 在 GitHub 查看
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsLinux
描述
Three paths ported from eddoursul's fork (F7) are compiled to `return false` on HIP:
- `head_mix_multi_enabled()` in `src/core/verify.cpp` (`STRATA_HEAD_MIX_MULTI`)
- the `head_mix_multi_on` gate in `MtpDrafter::record_rest` in `src/core/mtp.cpp`
- `one_token_self_commit()` in `src/core/verify.cpp` (`STRATA_ONE_TOKEN_COMMIT`)
That means nobody can try them on AMD without a rebuild. I patched them into **HIP opt-ins** (`=1` turns them on; the default stays off on HIP and CUDA is unchanged) and measured them on an R9700 (gfx1201):
| | decode tok/s (A/B/A/B) | answers |
|---|---|---|
| default | 91.6 / 91.6 | — |
| `STRATA_HEAD_MIX_MULTI=1` | **92.0 / 91.9 (+0.4%)**, faster on each of the 3 prompts in both rounds | bitwise identical |
| `STRATA_ONE_TOKEN_COMMIT=1` | 91.6 (0) | bitwise identical |
| both | 92.0 | bitwise identical |
Decode was measured with 3 fixed greedy prompts, 300 tokens and 3 rounds through the server, with a spread of 0.1-0.2% between server starts.
The gain is small, but it is exact and free. A default-on for gfx12 (or at least the opt-in) seems reasonable. Patch:
```diff
diff --git a/src/core/mtp.cpp b/src/core/mtp.cpp
index e73ce24..27ac0ea 100644
--- a/src/core/mtp.cpp
+++ b/src/core/mtp.cpp
@@ -852,7 +852,8 @@ bool MtpDrafter::record_rest(int step_row, cudaStream_t cs, std::string& err) {
}();
static const bool head_mix_multi_on = [] {
#if defined(STRATA_USE_HIP)
- return false;
+ const char* v = std::getenv("STRATA_HEAD_MIX_MULTI"); // HIP: opt-in (=1)
+ return v != nullptr && std::atoi(v) != 0;
#else
const char* v = std::getenv("STRATA_HEAD_MIX_MULTI");
return v == nullptr || std::atoi(v) != 0;
diff --git a/src/core/verify.cpp b/src/core/verify.cpp
index 6b80569..30ac512 100644
--- a/src/core/verify.cpp
+++ b/src/core/verify.cpp
@@ -126,7 +126,11 @@ bool sh_stream_on() {
// STRATA_HEAD_MIX_MULTI=0 does too): bitwise the same sums.
bool head_mix_multi_enabled() {
#if defined(STRATA_USE_HIP)
- return false;
+ static const bool on = [] {
+ const char* v = std::getenv("STRATA_HEAD_MIX_MULTI"); // HIP: opt-in (=1)
+ return v != nullptr && std::atoi(v) != 0;
+ }();
+ return on;
#else
static const bool on = [] {
const char* v = std::getenv("STRATA_HEAD_MIX_MULTI");
@@ -138,7 +142,11 @@ bool head_mix_multi_enabled() {
bool one_token_self_commit() {
#if defined(STRATA_USE_HIP)
- return false;
+ static const bool on = [] {
+ const char* v = std::getenv("STRATA_ONE_TOKEN_COMMIT"); // HIP: opt-in (=1)
+ return v != nullptr && std::atoi(v) != 0;
+ }();
+ return on;
#else
static const bool on = [] {
const char* v = std::getenv("STRATA_ONE_TOKEN_COMMIT");
```
**Not included:** the same treatment for `STRATA_MTP_SHARED_BRANCH` in `mtp.cpp` does not work on HIP. With `=1`, every request fails with `mtp: the shared expert's branch`, because the `cudaEventRecord(sh_fork_)` / `cudaStreamWaitEvent(side_)` pair fails, and `sh_fork_` / `side_` look like they are never created on HIP. That gate has to stay `false` unless the side stream is ported.
**Environment:** Strata v0.1.40.3 (d5ea713), Linux, AMD Radeon AI PRO R9700 (gfx1201, 32 GB), Ryzen 9 9900X, 30 GB RAM, TheRock ROCm 7.14.0a20260612 (`STRATA_ROCM_VERSION`), Coder IQ1_M, `--resident-experts`, `STRATA_HIP_WMMA=1`, `--prefill auto:32768`, all 12,288 experts in VRAM.
**Quality check used for every variant:** greedy needle retrieval at ~4K / ~16K / ~57K tokens plus a code task. All 4 pass, and the 4 answer hashes are identical to the baseline.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。