Issues / #703

#703 [AMD HIP] RX 6750 XT (gfx1031) works in heterogeneous layer split; long multimodal prompts can crash the engine

closed · @PJGV333 · 1 comentarios · En GitHub

BenchmarksSetup & installAMD / HIPModels & quantsLinux

Descripción

## Hardware

- Linux / CachyOS
- Ryzen 9 7950X
- 64 GB RAM
- GPU 0: RX 9070 XT 16 GB (`gfx1201`)
- GPU 1: RX 6750 XT 12 GB (`gfx1031`), connected through PCIe 4.0 x4

## Strata

- v0.1.38
- AMD HIP backend
- 262144 context
- Q4_0 KV
- Models tested:
  - Qwen3.8-Flash-Next IQ2_XS
  - Swift 1.5 IQ2_XS

`gfx1031` is currently marked as unvalidated. I added it locally to the HIP architecture whitelist and rebuilt Strata for:

```text
gfx1201;gfx1031
```

The RX 6750 XT works successfully together with the RX 9070 XT in heterogeneous `layer_split auto`.

## gfx1031 / dual-GPU results

### Long text generation

With Qwen3.8-Flash-Next IQ2_XS:

```text
--gpus 0,1
layer split auto
262,049 generated tokens
80.9 tok/s average
89.7% expert-cache hit
```

The full 262K generation completed without a GPU/backend crash.

This is especially encouraging because the RX 6750 XT is connected through only PCIe 4.0 x4.

### Swift 1.5 long generation

Swift 1.5 also works normally with the same heterogeneous split.

One long run completed:

```text
[strata] done: 75437 tokens in 1042 s (72.5 tok/s)
(stop, cancel=False), expert cache 87.5% hit
```

So long decode/generation itself appears stable on `gfx1201 + gfx1031`.

## Vision

CPU vision also works with the dual-GPU configuration in a fresh conversation.

Example:

```text
image prompt: ~344 tokens
901 tokens completed in 18 s
55.3 tok/s
90.5% expert-cache hit
```

The model correctly analyzed the image.

Short multimodal prompts therefore work with:

```text
gfx1201 + gfx1031
layer_split auto
CPU vision
```

## Problem: long-history prefill

Initially I thought the problem was related to vision because adding an image to an existing ~64.9K-token conversation caused the engine to exit during prompt processing.

A fresh image conversation worked immediately afterwards, and the same long multimodal prompt could be processed with the RX 9070 XT alone.

However, I have now reproduced the failure **without using an image**, so this does not appear to be vision-specific.

After Swift completed the 75,437-token generation shown above, I sent a simple text-only follow-up message:

> why did you stop?

Strata then had to prefill the existing ~75K-token conversation history and failed with:

```text
prefill PLE: native PLE postops launch:
no kernel image is available for execution on the device
```

This makes the current behavior look like:

```text
gfx1201 + gfx1031 long decode/generation     -> works
gfx1201 + gfx1031 short prefill              -> works
gfx1201 + gfx1031 fresh multimodal prompt    -> works
gfx1201 + gfx1031 long-history prefill       -> fails

gfx1201 alone + ~65K long multimodal prefill -> works
```

So the common factor now appears to be:

```text
large-history / long prefill
+
heterogeneous gfx1201 + gfx1031 layer split
+
native PLE postops
```

The exact HIP error:

```text
no kernel image is available for execution on the device
```

suggests that one of the native PLE post-op kernels used by the long-prefill path may not have a `gfx1031` kernel image in the multi-architecture build.

This would also explain why the previous 262K-token stress test worked: long **decode** is fine, while the failure appears when a subsequent request requires a large **prefill**.

## Summary

The RX 6750 XT (`gfx1031`) appears to work surprisingly well with Strata:

- heterogeneous `gfx1201 + gfx1031` layer split works
- long decode/generation works
- 262K stress generation completed successfully
- Swift 1.5 works
- CPU vision works
- fresh multimodal prompts work

The remaining issue appears to be specifically related to **long-history prefill / native PLE postops** on the heterogeneous split.

I can reproduce the issue, provide full logs, test patches, or run additional A/B tests if useful.

Thanks for the excellent work on Strata. The performance on this heterogeneous AMD setup has been surprisingly good, especially considering that the RX 6750 XT is running through only a PCIe 4.0 x4 link.

En el sitio

Enlaces a install, modelos, releases.