Pull requests / #1579
#1579 bench: record dual-GPU auto-layout measurements for #1352
open · @Momoyeyu · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsLinux
Description
## Summary Related to #1352; this intentionally does not close it. Adds a community measurement record under `bench/results/2026-10-08-community-2x-rtx-4080-super/`, without changing the engine or server. - Pinned main `d5ea713` and Hardin22/Strata-DualGPU `03a5ec9`, the same verified IQ3_S model and frozen prompt IDs, both physical GPU orders, four requested pipeline/async settings, and three repetitions. - Two RTX 4080 SUPER devices reporting 32,760 MiB each, Linux, NVIDIA driver 610.43.03. This is not the reporter's two 16 GB cards or original configuration. - Preserve all 48 engine logs, manifests, 48 warmups and 96 measured requests, output IDs, timing distributions, configuration/input provenance, a dry-run-by-default runner and CPU tests. All measured requests returned 512 tokens without prompt reuse. The useful finding is that identical requested flags do not establish identical paths: main selected split 24 and held all experts in VRAM, so its requested async rows did not start an adaptive worker. The fork selected different splits/cache residency depending on the setting and order. Pipeline activation was checked in the logs. Only 12 of 48 matched main/fork outputs were token-identical; some fork pipeline cells also varied across repetitions. The report therefore presents configuration observations, not identical-output speedup ratios, an isolated async benefit, a quality result, or a disproof of the original slowdown. It explicitly documents the fixed engine/order sequence, warmup and page-cache limitations. It does not reproduce #1468's large prefill chunks. Private workspace prefixes and GPU UUIDs are redacted. `artifact-hashes.json` covers the published bytes; internal provenance hashes refer to the original capture before redaction. #### Test plan - [x] Completed the four real dual-GPU matrices sequentially within a bounded execution window. - [x] Checked all expected request/case counts, pinned engine/source identities, input hashes, generated lengths, no-reuse fields and DONE timings against the raw logs. - [x] Audited effective split/cache/pipeline/async startup behavior and retained raw output IDs. - [x] `PYTHONDONTWRITEBYTECODE=1 python3 -m unittest -v test_run_experiments` — 9 CPU/mock runner tests passed; these are not inference measurements. - [x] JSON/JSONL parsing, all 71 published artifact checksums, sanitized-copy verification and sensitive-string audit. - [x] `git diff --check`; merges cleanly with current upstream main, without rebasing the pinned measurements.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.