贡献 / #1716

#1716 bench: community report — Strix Halo laptop (Radeon 8060S, 70 W), v0.1.41, UD-IQ4_XS

open · @nsbradley88 · 0 评论 · 去 GitHub 看

BenchmarksSetup & installAMD / HIPDocumentation

说明

Results-only community report (no engine changes), following `docs/COMMUNITY_BENCHMARKS.md`.

- **Hardware:** HP ZBook Ultra G1a laptop, Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), 128 GB unified memory, **70 W package limit** under load.
- **Software:** Strata v0.1.41 (`fb58e0d`) source build, ROCm 7.14.1, table `100401`; `docs/STRIX_HALO.md` §5 fast config plus `--spec 5 --mtp-window 8192`.
- **Model:** Unsloth UD-IQ4_XS, all 24,576 experts on the GPU.
- **Results** (medians of 4–6 fresh-process runs at 8,192 / 32,768 / 131,072 prompt tokens): prompt 969 / 1,067 / 1,031 tok/s; decode 59.2 / 57.8 / 42.0 tok/s. Needle 15/15, agentic coding suite 5/5.
- **Limitations:** prompt throughput is ~20% below the §6 desktop table, which we attribute to the 70 W limit (config, pack, ROCm runtime and kernel args were A/B'd); TTFT not measured separately; `tools/needle_bench.py` not run.
- Also includes same-session comparisons (`--spec 4/5/6`, the §6 flags, ROCm 10.2 vs 7.14.1, batch slots) and a failure note: setup's `--expert-cache auto` caused a global OOM here (#1715).

Write-up and configs: https://github.com/nsbradley88/halo-tuning

🤖 Generated with [Claude Code](https://claude.com/claude-code)

本站相关内容

相关页面的快捷入口。