Pull requests / #1435

#1435 docs: report RTX 8000 measurements of existing runtime switches

open · @agorevski · 0 comments · View on GitHub

BenchmarksNVIDIA / CUDAModels & quantsDocumentation

Description

## Title
<!-- If Applicable, reference the GitHub issue -->
Document measured RTX 8000 results for existing runtime switches

Issue: Not linked to an existing issue.

## Summary
<!-- Quick Summary of changes -->
Document the already-existing `STRATA_BF16_TC` and `STRATA_SELECT_SIMT` switches on a Quadro RTX 8000, including their architecture limits, measured prompt/decode trade-offs and rounding caveats. This helps owners choose settings from measured results instead of assuming a universal speedup.

**No new kernel, engine code, default change or dependency on another optimization PR.**

## What changed
<!-- Specifics on files changed, and what changes were made there -->
- Only `docs/OLDER_GPUS.md`: add an RTX 8000 existing-switch subsection and links to the baseline implementations.
- Explain sm_75's FP16 tensor cores and lack of native BF16/TF32, opt-in config usage, and how to turn the switches off.
- Report three-arm median prompt/decode rates at 537, 4,057 and 32,057 prompt tokens with the machine, power limit, model, memory settings and method.
- Distinguish prompt gains from decode results; explain changed replies and the limits of five arithmetic/JSON smoke checks and a synthetic QSA FP64/selected-ID check.
- Record negative tuning findings: doubling the prefill ceiling consumed another 3 GiB of prompt loan for about 0.6% more 32K prompt throughput with only 24 MiB free at the final graph capture; CPU sharing did not improve the two tested prompt sizes.

## Extra Notes
<!-- Any extra notes, delete if there are none -->
### Existing measurement evidence

The measurements were collected on 2026-10-07 on physical GPU 3: a 48 GiB Quadro RTX 8000 at a **260 W** limit, Xeon W-2295, CUDA 12.4 and GCC 13. They are not measurements at the launcher's 200 W setting and are not promised on every sm_75 card.

The measured source-built engine used the v0.1.40.3 baseline with other local changes integrated; unrelated opt-in changes were disabled in the selected arms. **This isolated documentation branch did not produce or reproduce those runtime numbers.**

Previously collected evidence, retained outside the commit in the session's `files/` directory:
- `bench-tc0-result.json`, `bench-tc1-result.json`, `bench-simt-result.json`: three fresh seeded prompts per size, 128 output tokens per request, uncached prompts; median table values and 5/5 basic checks per arm. BF16_TC changed reply hashes in 3/9 paired requests.
- `bench-prefill16k-result.json`, `bench-cpu-share-result.json`: tuning medians and 5/5 basic checks per arm. CPU sharing was tested only at 537 and 4,057 tokens.
- `rtx8000-runtime-benchmark.log`, `rtx8000-runtime-tuning.log` (the persisted tuning log), corresponding `bench-*-engine.log` files, `benchmark_rtx8000.py` and `bench-*-config.json`: cross-check rate samples, switch settings, request method, engine arguments and memory headroom.
- `rtx8000-scorer-bench.log`: the separate synthetic 131,072-context QSA check passed its FP64 gate and matched selected IDs for 256/256 queries. This is not an arbitrary-prompt parity or model-quality claim.

### Validation performed for this documentation change

Run from `files/pr-worktrees/rtx8000-measurements`:
- `python ../../validate_rtx8000_measurements.py` — **passed**. Cross-checked all five persisted result sets, raw-sample medians against summaries/logs, table rounding, three requests per size, output/cache counts, 5/5 checks per arm, changed reply hashes, engine arguments and memory values, and the QSA evidence. Also verified the switches already exist at base `d5ea7133741e67743c0e886bb426c0ce8d69cf6c`, both new relative links resolve, and only `docs/OLDER_GPUS.md` differs from that base. This evidence-check script is a session artifact, not part of the PR.
- `git diff --check` — **passed**, no whitespace errors.
- `git diff d5ea7133741e67743c0e886bb426c0ce8d69cf6c HEAD --check` — **passed**, no whitespace errors in the committed change.
- `git diff d5ea7133741e67743c0e886bb426c0ce8d69cf6c HEAD --name-only` — **passed**, only `docs/OLDER_GPUS.md`.

No GPU tests, benchmark reruns or builds were run for this documentation-only change.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.