Pull requests / #758
#758 bench: community MI50 (gfx906) — Strata 0.1.38 speed and recall
closed · @xxDoman · 0 Kommentare · Auf GitHub
BenchmarksAMD / HIPModels & quantsDocumentation
Beschreibung
Community benchmark report for Strata 0.1.38 on an **AMD MI50 32GB (Vega 20, gfx906)** — an experimental, unsupported architecture. **Hardware:** MI50 32GB, Intel i5-12400F (12 threads), 64 GB RAM, ROCm 7.2.1, Ubuntu 26.04, PCIe Gen4 x16. gfx906 host build packaged as a Docker image. **Model / config:** `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF` (rev `ed59f92082b1e93c0e96d60a8b11aab089b52f09`) IQ2_XS, both GGUF SHA-256 hashes verified against the published LFS hashes. Context 32,768; INT8 KV; expert cache `auto` (19,427 slots, 26.10 GiB); prefill 2048; MTP `--spec 3 --spec-min-p 0.9`; vision off. **Results (median of 3 runs, 256-token cap, zero reused tokens):** | Prompt tokens | Prompt tok/s | Decode tok/s | TTFT s | | ---: | ---: | ---: | ---: | | 4,096 | 330.8 | 36.4 | 12.42 | | 24,576 | 329.5 | 35.0 | 74.65 | Recall: `tools/needle_bench.py --lengths 8k,24k --depths 10,50,90` → **6/6 found**. Peak host RAM 38.5 GiB, peak swap ~8.1 GiB; MI50 filled to ~32 GB. **Notes:** 32,768 was used as the long length instead of 32,768-prompt because prompt + 256 output exceeds the 32,768 context (requests are never truncated; a 32,768-token prompt returns HTTP 400). The 128,000 length does not fit and was not run. Harness is the RTX-5090 community script adapted only for the container port/pack and a required API-key header. Full details in the report README.
Mehr auf der Site
Links zu Install, Modellen, Releases.