Pull requests / #677

#677 bench: Linux + 2x AMD Instinct MI50 16 GB (gfx906), Coder IQ1_M, 4K to 128K prompt tokens

closed · @JeanP00l · 0 comentários · No GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

Descrição

Community benchmark per `docs/COMMUNITY_BENCHMARKS.md`, in the format of the RTX 5090 report, using its `benchmark.py` unchanged.

**Hardware:** 2x AMD Instinct MI50 16 GB (gfx906, 85 W each, PCIe 3.0 x16), Xeon E5-2666 v3 (AVX2), 32 GB DDR4.

**Build:** the gfx906 build of #638, with #639 and #640 merged (branch `mi50-bench-build` at `ea51fd3`, based on `main` 99f3dbd).

**Configuration:**
- Coder IQ1_M, 131,072-token context, `--kv int8 --kv-resident 32768`;
- layer split 27, all 12,288 experts in VRAM;
- MTP `--spec 4`, `STRATA_ARENA_MMAP=1`.

| Prompt tokens | Reused | Prompt tok/s | Decode tok/s | TTFT s |
| ---: | ---: | ---: | ---: | ---: |
| 4,096 | 0 | 320.9 [309.9-324.7] | 50.1 [49.3-54.0] | 12.81 |
| 32,768 | 0 | 522.8 [522.6-523.4] | 47.8 [47.6-48.5] | 62.81 |
| 128,000 | 0 | 516.8 [516.5-516.9] | 45.7 [44.8-51.5] | 248.09 |

**Recall:** `needle_bench` found 6 of 6 needles at 32K and 128K.

**Memory:**
- VRAM peak: 15.4 / 14.7 GiB.
- `MemAvailable` never below 25.1 GiB of 31.2.

The folder holds the per-run results with output text, the engine log, a telemetry log sampled once a second, and the server config. The README lists the limits (one machine, synthetic workload, no controlled comparison).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.