Pull requests / #1579

#1579 bench: record dual-GPU auto-layout measurements for #1352

open · @Momoyeyu · 0 comentarios · En GitHub

BenchmarksNVIDIA / CUDAModels & quantsLinux

Descripción

## Summary

Related to #1352; this intentionally does not close it. Adds a community measurement record under `bench/results/2026-10-08-community-2x-rtx-4080-super/`, without changing the engine or server.

- Pinned main `d5ea713` and Hardin22/Strata-DualGPU `03a5ec9`, the same verified IQ3_S model and frozen prompt IDs, both physical GPU orders, four requested pipeline/async settings, and three repetitions.
- Two RTX 4080 SUPER devices reporting 32,760 MiB each, Linux, NVIDIA driver 610.43.03. This is not the reporter's two 16 GB cards or original configuration.
- Preserve all 48 engine logs, manifests, 48 warmups and 96 measured requests, output IDs, timing distributions, configuration/input provenance, a dry-run-by-default runner and CPU tests. All measured requests returned 512 tokens without prompt reuse.

The useful finding is that identical requested flags do not establish identical paths: main selected split 24 and held all experts in VRAM, so its requested async rows did not start an adaptive worker. The fork selected different splits/cache residency depending on the setting and order. Pipeline activation was checked in the logs.

Only 12 of 48 matched main/fork outputs were token-identical; some fork pipeline cells also varied across repetitions. The report therefore presents configuration observations, not identical-output speedup ratios, an isolated async benefit, a quality result, or a disproof of the original slowdown. It explicitly documents the fixed engine/order sequence, warmup and page-cache limitations. It does not reproduce #1468's large prefill chunks.

Private workspace prefixes and GPU UUIDs are redacted. `artifact-hashes.json` covers the published bytes; internal provenance hashes refer to the original capture before redaction.

#### Test plan

- [x] Completed the four real dual-GPU matrices sequentially within a bounded execution window.
- [x] Checked all expected request/case counts, pinned engine/source identities, input hashes, generated lengths, no-reuse fields and DONE timings against the raw logs.
- [x] Audited effective split/cache/pipeline/async startup behavior and retained raw output IDs.
- [x] `PYTHONDONTWRITEBYTECODE=1 python3 -m unittest -v test_run_experiments` — 9 CPU/mock runner tests passed; these are not inference measurements.
- [x] JSON/JSONL parsing, all 71 published artifact checksums, sanitized-copy verification and sensitive-string audit.
- [x] `git diff --check`; merges cleanly with current upstream main, without rebasing the pinned measurements.

En el sitio

Enlaces a install, modelos, releases.