Pull requests / #1459

#1459 bench: RX 7900 XTX + 6800 XT helper layout, 1M and long-output results

open · @saikiran-rs · 0 Kommentare · Auf GitHub

BenchmarksDocumentationLinux

Beschreibung

Results-only community report for a Linux RX 7900 XTX + RX 6800 XT helper layout with a Xeon E5-2696 v4 and 121 GiB OS-visible RAM.

## Measurements

- Three 8192-token coding generations: 97.2 / 96.7 / 97.0 tok/s. Long prose: 65.0 / 63.9, with actual output lengths retained. The original 100TG gate failed.
- Full usable1M request: 1048064 input +504 generated +8 API-reserved tokens; all three recall markers found. PP591.7 vs587.4 baseline, full-depth decode39.2 vs39.3.
- Matched first200 HellaSwag/WinoGrande sample: baseline181/173, final181/175, zero unparseable answers/errors.

## Reproduction and limits

Includes exact compressed prompts, per-request engine/client timings, actual token counts, portable configs, pinned model URLs/hashes, source/build provenance and a standalone client/quality harness. Historical source revisions are preserved on the contributor's fork. Offline fixture hashes and pinned-API token counts were checked; the SSE parser was tested with a synthetic stream.

These are historical measurements from a combined0.1.40.2-derived research build, not isolated measurements of a new PR or stock0.1.40.3. All changed settings and inherited contributor work are identified. Q8-kernel-only whole-model results include regressions. The separately managed NVIDIA service's host load was not isolated. No new inference was run for publication.

This report follows docs/COMMUNITY_BENCHMARKS.md and is separate from the helper-allowance engine proposal. No universal speedup or general quality equivalence is claimed.

Mehr auf der Site

Links zu Install, Modellen, Releases.