Pull requests / #1437

#1437 perf(cuda): add opt-in shape-tuned sm75 interleaved verify

open · @agorevski · 0 评论 · 在 GitHub 查看

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

描述

## Title
<!-- If Applicable, reference the GitHub issue -->
Opt in to measured sm_75 interleaved dense verify projections

Issue: Not linked to an existing issue.

## Summary
<!-- Quick Summary of changes -->
Allow Turing (sm_75) to reuse interleaved Q8_1 activations for selected 2-4-token dense verify projections. The existing interleaved kernels preserve the exact multi-column output while reducing repeated activation reads. Selection is limited to measured IQ4_XS/Q4_K/Q5_K/Q6_K shapes and requires exactly `STRATA_MMVQ_IL=1`; unset, `0`, and `true` leave sm_75 on its existing path.

All sm_75 devices are eligible, not only the measuring card. However, tuning and hardware validation are **Quadro RTX 8000 only**; performance on other Turing cards is not established. The sm_80+ automatic table remains unchanged. Pascal, Volta and HIP do not gain an interleaved path.

## What changed
<!-- Specifics on files changed, and what changes were made there -->
- `src/kernels/cuda/native_mmvq.cu`: add the strict opt-in sm_75 shape/window table, exact reduction-width eligibility, validated test tuning overrides, and thread-local last-launch path observation (no device work or synchronization; works during graph capture).
- `include/strata/kernels/native_mmvq.hpp`: expose reduction width in the eligibility query and last-path observation.
- `src/core/verify.cpp`: query with the actual reduction width before interleaving dense projections and the output head; unlisted shapes retain the existing fallback.
- `src/kernels/mmvq_il_parity.cpp`: cover dense model shapes and row/reduction tails, assert real dispatch rather than accidentally timing fallback, compare bitwise outputs for table and forced 1/2/4-row paths, and add `--sm75-bench` with cold-weight rotation and seven alternating timing pairs including pack cost.
- `docs/OLDER_GPUS.md`: document opt-in/default safety, architecture eligibility, RTX 8000-only measurements, and parity/benchmark usage.

No BF16 cache, peer placement, CPU changes, model/source configs, logs, or installed engines are included.

## Extra Notes
<!-- Any extra notes, delete if there are none -->
### Existing measurement evidence
The table keeps 46/69 measured shape/window cells: every paired trial in both seven-pair sweeps saved at least 3%, charging the full interleave launch to each MMVQ. Cross-checking the retained cells against `build-sm75-tune/sm75-bench.log` and `sm75-bench-repeat.log` passed.

Previously recorded controlled whole-model decode medians (three runs per prompt, 128 outputs each) were:

| Prompt tokens | baseline (tok/s) | `STRATA_MMVQ_IL=1` (tok/s) |
| --- | ---: | ---: |
| 537 | 83.7 | 88.9 |
| 4,057 | 75.4 | 80.4 |
| 32,057 | 68.4 | 72.1 |

Quadro RTX 8000 (physical GPU 3), 260 W, Xeon W-2295, CUDA 12.4/GCC 13, IQ3_S + MTP, spec 4, fixed 19,000 expert slots, automatic prefill 8,192, INT8 KV, maximum context 262,144 / resident 32,768, PCIe fraction 0.20. Prompt cache and adaptation were disabled. The baseline already had `STRATA_BF16_TC=1`; BF16 cache, SIMT scorer and CPU prefill sharing were off in both variants. The recorded `bench-tc1-result.json` and `bench-interleaved-result.json` medians were independently checked. These measurements are prior evidence, **not new isolated-branch whole-model runs**.

### Isolated branch validation
Run from the independent `perf/sm75-interleaved-verify` worktree based on `d5ea7133`. Fresh objects, not reused objects from the dirty checkout; only the existing llama.cpp vendor source is shared read-only.

```sh
PATH=/usr/bin:/bin:/home/algore/miniconda3/bin cmake -S . -B build-pr-sm75 \
  -DSTRATA_ENABLE_CUDA=ON -DSTRATA_BUILD_TESTS=ON -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_CUDA_ARCHITECTURES=75 \
  -DCMAKE_CUDA_COMPILER=/home/algore/miniconda3/bin/nvcc \
  -DCMAKE_C_COMPILER=/usr/bin/gcc -DCMAKE_CXX_COMPILER=/usr/bin/g++ \
  -DCMAKE_CUDA_HOST_COMPILER=/usr/bin/g++ \
  -DSTRATA_GGML_DIR=/home/algore/GIT/strata/third_party/llama.cpp
PATH=/usr/bin:/bin:/home/algore/miniconda3/bin cmake --build build-pr-sm75 --target strata mmvq_il_parity -j2
```

Results: configuration and fresh builds of `strata` and `mmvq_il_parity` passed (exit 0), using CUDA 12.4 and GCC 13. Compiler warnings were emitted; no build errors.

All GPU commands acquire the same exclusive lock for their full lifetime and expose only physical GPU 3:

```sh
flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 STRATA_MMVQ_IL=1 ./build-pr-sm75/mmvq_il_parity
flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 STRATA_MMVQ_IL=0 ./build-pr-sm75/mmvq_il_parity
flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 STRATA_MMVQ_IL=true ./build-pr-sm75/mmvq_il_parity
flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env -u STRATA_MMVQ_IL CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 ./build-pr-sm75/mmvq_il_parity
git diff --check
```

Results: all four GPU parity commands passed (exit 0), bitwise across every case and 2-4-token window. Each policy actually exercised 387 forced interleaved launches. With `=1`, the automatic table chose 67 interleaved / 62 fallback cases (the harness includes overlapping case sets); with `=0`, `=true`, and unset it chose 0 interleaved / 129 fallback cases each. Exact-shape and strict opt-in policy assertions passed. `git diff --check` passed.

Other sm_75 hardware, sm_80+, Pascal/Volta and HIP were not hardware-tested or compiled for this PR. No new whole-model server run or performance sweep was performed in this worktree; no power limits were changed. The opt-in table is not a claim of speedup on untested cards.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。