Issues / #1375

#1375 Strata on an RX 6600 XT (gfx1032, 8 GB) and a Xeon E5-2678 v3 with DDR3

open · @Stataltex · 0 comentários · No GitHub

BenchmarksSetup & installAMD / HIPModels & quantsDocumentation

Descrição

Hello,

This is a short report on running Strata on an older Xeon / DDR3 machine, with an 8 GB mid-range card that is not in the supported list. It works well. Here are my results and settings. They may help other users with similar hardware. Thank you for the work on this project.

The setup was done with the help of Claude (Opus 5.5) and Codex (GPT-6.1 Sol).

## Machine

| | |
|---|---|
| Board | Huananzhi X99-TF |
| CPU | Intel Xeon E5-2678 v3 (Haswell-EP), 12 cores / 24 threads, 3.3 GHz on all cores with a BIOS turbo unlock patch, AVX2, no AVX-512 |
| RAM | 4 x 16 GB DDR3-1600 registered ECC, quad channel, 51.2 GB/s |
| GPU | AMD Radeon RX 6600 XT, 8 GB, gfx1032 |
| PCIe | 3.0 x8, limited by the X99 platform, 6.7 GB/s measured by Strata |
| Storage | NVMe PCIe 3.0 SSD for the model files |
| OS | Debian 13, kernel 7.1.8 |
| Strata | v0.1.40.1 (82f46a8), HIP engine built on the machine for gfx1030, `-DSTRATA_PREFILL_MMQ=ON` |
| ROCm | 7.14.0a20260612 from the `gfx103X-all` wheels index |

Models:

- GSQ-RCO IQ3_XXS: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
- Uncensored IQ3_XXS, compat-bf16 pack: https://huggingface.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF

## Installation

`setup.py` and `cmake/hip_backend.cmake` do not accept gfx1032, and `HSA_OVERRIDE_GFX_VERSION` does not change what KFD reports to the setup. So I used a small local wrapper of `setup.py` that reports gfx1032 as gfx1030, and I built the engine for gfx1030. I run it with `HSA_OVERRIDE_GFX_VERSION=10.3.0`. `strata-device --selftest` passes.

The pinned ROCm version, 7.10.0a20251120, was not available from the `gfx103X-all` index, so I used `STRATA_ROCM_VERSION=7.14.0a20260612`.

## Settings

I tested the performance settings with the same French prompts. The read speeds are for a 10,567-token prompt. Most settings were tested once per value. The VRAM reserve was tested three times per value.

| Setting | Selected value | Comparison or reason |
|---|---|---|
| Prompt chunks | `--prefill auto` | Uncensored: about 318 tok/s with auto-selected 3840-token chunks, against 79 at 512 (the value in docs/ORCA.md). GSQ-RCO: about 250 tok/s at auto-selected 2304, against 133 when requesting 3840, which the engine reduced to 960. |
| Prompt FP16 | `STRATA_HIP_PROMPT_F16=1` | Uncensored: 318 tok/s enabled, against 229 disabled. |
| Expert copies | `STRATA_HIP_ADAPT_KERNEL_COPY=1` | One GPU hang with an SDMA page fault during a long prompt before, none after. |
| Context | `--max-context 131072 --kv int8 --kv-resident 20480` | GSQ-RCO: about 252 tok/s at 128K, against 256 with the initial 32K configuration and 158 at 256K. The 32K baseline also used the default draft vocabulary and fully resident KV. |
| Draft vocabulary | `--mtp-draft-vocab data/draft_vocab_fr.bin` | Most of my prompts are in French. The draft head uses 77.6 MiB, against 178.4 MiB with the default vocabulary. |
| VRAM reserve | `--vram-reserve-mib 700` | Uncensored: 326 to 327 tok/s, against 317 at 500 MiB. The lower reserve added 89 expert slots but gave no clear decode gain. |

`STRATA_DENSE_MMQ`, `STRATA_PREFILL_RING`, `--pcie-frac` and `--pool-workers` were also tried. They did not show a consistent gain, so they stay at their defaults.

Both models use the original MTP runtime with `--spec 4 --spec-min-p 0.5`, as described in docs/ORCA.md. Vision uses `--vision` and the CPU encoder with 12 threads, `min_tokens 512` and `max_tokens 2048`.

## Results

| | GSQ-RCO IQ3_XXS | Uncensored IQ3_XXS |
|---|---|---|
| Expert cache in VRAM | 1134 slots, 1.87 GiB | 1236 slots, 2.50 GiB |
| Decode, short answer | 16 to 17 tok/s | 14 to 16 tok/s |
| Prompt read, 10,567 tokens | about 250 tok/s | about 320 tok/s |
| Decode after this prompt | about 17 tok/s | about 17.5 tok/s |
| Prompt read, 42,176 tokens | | 341 tok/s |
| MTP drafts accepted, long prompt | 65 to 72 % | 63 to 71 % |
| Server cold start | | 46 s |
| System RAM used after loading | about 46 GiB | about 56 GiB |

- GSQ-RCO found a sentence placed in the middle of a 115,350-token prompt, with the 128K configuration.
- Vision with the CPU encoder: an 800 x 600 image takes about 8 to 9 s to encode. A 1920 x 1080 screenshot of dense text takes about 65 s, and the requested line was copied correctly.

No site

Links install, modelos, releases.