Pull requests / #1596

#1596 bench: MI50 32 GB (gfx906) on 0.1.40.1, 252K needles, a 16 GB-limit run, temperatures; 0.1.40.3 / 0.1.41 retest

open · @mathcuei · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Descripción

Follow-up to #1463 (the review there asked for a `bench/results` folder and for the 16 GB-limit 252K case to be repeated on 0.1.40.3).

## What changed
- New folder `bench/results/2026-10-06-community-mi50-32gb/`: README, configs, engine logs, harness scripts and outputs, temperature readings (no file over 100 KB; paths replaced by `<strata-dir>`, `<data-dir>`, `<models-dir>`).
- `retest-2026-10-08/` inside it: the repeat of the case that answered `!!!!...` on 0.1.40.1 (MI50 32 GB limited to 16 GB, `--max-context 262144`, ~250K-token prompt, three needles).
- One row in `bench/results/COMMUNITY.md`.

## Retest result
On **0.1.40.3** (`d5ea713`) and **0.1.41** (`fb58e0d`), two different prompts each: **3 of 3 needles in all 4 runs, no `!!!!`** in the answers or the logs. Prefill 361-364 tok/s cold at 252K, decode 29-34 tok/s afterwards. 0.1.41 built for gfx906 with no local patch; 0.1.40.3 needed only the one-line `cudaEventBlockingSync` define from #1396.

## Extra Notes
- Junction temperature peaked at 106 C on both versions (about 8 minutes at 105 C or above, the card's own critical threshold), so the card was throttling itself. This is stated in the folder's README; it limits how far the result can be read.
- It does not say what caused the original `!!!!`: 0.1.40.1 itself was not rerun, and that failure was a single run.
- Runs were driven by Claude Code on the submitter's own machine.

En el sitio

Enlaces a install, modelos, releases.