Pull requests / #1717

#1717 bench: community report, Radeon 8060S (gfx1151) on a Ryzen AI Max+ 395, UD-IQ4_XS, engine 0.1.41

open · @libratechw · 0 commentaires · Sur GitHub

BenchmarksAMD / HIPDocumentation

Description

Issue: Related to #1356; this report records the retest and does not request closing the issue.

## Summary

Add measurements on an EVO-X2 with Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), 128 GB UMA, using an unmodified v0.1.41 HIP build and pinned Unsloth UD-IQ4_XS weights. This is a second gfx1151 report next to `2026-10-05-community-gfx1151` (#917), which measured the hipBLASLt table, dense MMQ and the shared-expert stream on 0.1.39 plus unmerged PRs; this one measures the 0.1.41 defaults against the fast configuration.

The default configuration scores 65/69 in both repetitions of the existing P40 correctness suite, with identical generated content on all 69 tasks. On the same-day matched 128,000-token synthetic-Python workload, the fast configuration (all nine switches of `docs/STRIX_HALO.md` section 5) measures 1,108.6 prompt tok/s and 47.5 decode tok/s, compared with 617.4 and 42.1 on the default configuration. These are below the 1,320 / 51.4 and 1,299 / 50.7 in section 6; the README lists the differing conditions and says the cause was not isolated. The official recall probe finds all nine code words at actual inputs up to 251,817 tokens.

#1356 reports a hang and garbled output when two long prompts run in parallel. This report tries to reproduce that with public synthetic prompts of about the same length (two simultaneous ~19,600-token prompts, QFUSE on/off × prefill-borrow on/off). All eight requests complete. The hang and garbling were not reproduced on these inputs, with one paired trial per cell; this does not show that #1356 is fixed.

## What changed

- Add `bench/results/2026-10-09-community-gfx1151-ud-iq4xs-0.1.41/`: README, reproduction commands, summary JSON and CSVs, server configurations, two engine logs and provenance. `TRIMMED.md` lists what was left out.
- Add small report-local scripts that reuse the public P40/V100 workloads and the official needle and batch tests. No existing task, grader or engine source changes.
- Add one row to `bench/results/COMMUNITY.md`.

## Limitations

- Fast-arm correctness, a real agent harness, long-session soak, vision and other backends were not tested. The correctness score applies to the default arm only.
- The P40 report uses different weights, hardware and engine version, so this is not a controlled cross-model quality comparison.
- The fast arm's 16,384-token prefill lends 2,939 cache slots to the prompt path (engine warning in `raw/fast-long-engine.log`); no arm avoiding this was measured.

Sur le site

Liens install, modèles, releases.