Pull requests / #1437
#1437 perf(cuda): add opt-in shape-tuned sm75 interleaved verify
open · @agorevski · 0 commentaires · Sur GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
## Title <!-- If Applicable, reference the GitHub issue --> Opt in to measured sm_75 interleaved dense verify projections Issue: Not linked to an existing issue. ## Summary <!-- Quick Summary of changes --> Allow Turing (sm_75) to reuse interleaved Q8_1 activations for selected 2-4-token dense verify projections. The existing interleaved kernels preserve the exact multi-column output while reducing repeated activation reads. Selection is limited to measured IQ4_XS/Q4_K/Q5_K/Q6_K shapes and requires exactly `STRATA_MMVQ_IL=1`; unset, `0`, and `true` leave sm_75 on its existing path. All sm_75 devices are eligible, not only the measuring card. However, tuning and hardware validation are **Quadro RTX 8000 only**; performance on other Turing cards is not established. The sm_80+ automatic table remains unchanged. Pascal, Volta and HIP do not gain an interleaved path. ## What changed <!-- Specifics on files changed, and what changes were made there --> - `src/kernels/cuda/native_mmvq.cu`: add the strict opt-in sm_75 shape/window table, exact reduction-width eligibility, validated test tuning overrides, and thread-local last-launch path observation (no device work or synchronization; works during graph capture). - `include/strata/kernels/native_mmvq.hpp`: expose reduction width in the eligibility query and last-path observation. - `src/core/verify.cpp`: query with the actual reduction width before interleaving dense projections and the output head; unlisted shapes retain the existing fallback. - `src/kernels/mmvq_il_parity.cpp`: cover dense model shapes and row/reduction tails, assert real dispatch rather than accidentally timing fallback, compare bitwise outputs for table and forced 1/2/4-row paths, and add `--sm75-bench` with cold-weight rotation and seven alternating timing pairs including pack cost. - `docs/OLDER_GPUS.md`: document opt-in/default safety, architecture eligibility, RTX 8000-only measurements, and parity/benchmark usage. No BF16 cache, peer placement, CPU changes, model/source configs, logs, or installed engines are included. ## Extra Notes <!-- Any extra notes, delete if there are none --> ### Existing measurement evidence The table keeps 46/69 measured shape/window cells: every paired trial in both seven-pair sweeps saved at least 3%, charging the full interleave launch to each MMVQ. Cross-checking the retained cells against `build-sm75-tune/sm75-bench.log` and `sm75-bench-repeat.log` passed. Previously recorded controlled whole-model decode medians (three runs per prompt, 128 outputs each) were: | Prompt tokens | baseline (tok/s) | `STRATA_MMVQ_IL=1` (tok/s) | | --- | ---: | ---: | | 537 | 83.7 | 88.9 | | 4,057 | 75.4 | 80.4 | | 32,057 | 68.4 | 72.1 | Quadro RTX 8000 (physical GPU 3), 260 W, Xeon W-2295, CUDA 12.4/GCC 13, IQ3_S + MTP, spec 4, fixed 19,000 expert slots, automatic prefill 8,192, INT8 KV, maximum context 262,144 / resident 32,768, PCIe fraction 0.20. Prompt cache and adaptation were disabled. The baseline already had `STRATA_BF16_TC=1`; BF16 cache, SIMT scorer and CPU prefill sharing were off in both variants. The recorded `bench-tc1-result.json` and `bench-interleaved-result.json` medians were independently checked. These measurements are prior evidence, **not new isolated-branch whole-model runs**. ### Isolated branch validation Run from the independent `perf/sm75-interleaved-verify` worktree based on `d5ea7133`. Fresh objects, not reused objects from the dirty checkout; only the existing llama.cpp vendor source is shared read-only. ```sh PATH=/usr/bin:/bin:/home/algore/miniconda3/bin cmake -S . -B build-pr-sm75 \ -DSTRATA_ENABLE_CUDA=ON -DSTRATA_BUILD_TESTS=ON -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_CUDA_ARCHITECTURES=75 \ -DCMAKE_CUDA_COMPILER=/home/algore/miniconda3/bin/nvcc \ -DCMAKE_C_COMPILER=/usr/bin/gcc -DCMAKE_CXX_COMPILER=/usr/bin/g++ \ -DCMAKE_CUDA_HOST_COMPILER=/usr/bin/g++ \ -DSTRATA_GGML_DIR=/home/algore/GIT/strata/third_party/llama.cpp PATH=/usr/bin:/bin:/home/algore/miniconda3/bin cmake --build build-pr-sm75 --target strata mmvq_il_parity -j2 ``` Results: configuration and fresh builds of `strata` and `mmvq_il_parity` passed (exit 0), using CUDA 12.4 and GCC 13. Compiler warnings were emitted; no build errors. All GPU commands acquire the same exclusive lock for their full lifetime and expose only physical GPU 3: ```sh flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 STRATA_MMVQ_IL=1 ./build-pr-sm75/mmvq_il_parity flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 STRATA_MMVQ_IL=0 ./build-pr-sm75/mmvq_il_parity flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 STRATA_MMVQ_IL=true ./build-pr-sm75/mmvq_il_parity flock -x /home/algore/.copilot/session-state/a7c2bcbb-4264-4eb2-879d-c71a90285b0b/files/pr-gpu3.lock env -u STRATA_MMVQ_IL CUDA_DEVICE_ORDER=PCI_BUS_ID CUDA_VISIBLE_DEVICES=3 ./build-pr-sm75/mmvq_il_parity git diff --check ``` Results: all four GPU parity commands passed (exit 0), bitwise across every case and 2-4-token window. Each policy actually exercised 387 forced interleaved launches. With `=1`, the automatic table chose 67 interleaved / 62 fallback cases (the harness includes overlapping case sets); with `=0`, `=true`, and unset it chose 0 interleaved / 129 fallback cases each. Exact-shape and strict opt-in policy assertions passed. `git diff --check` passed. Other sm_75 hardware, sm_80+, Pascal/Volta and HIP were not hardware-tested or compiled for this PR. No new whole-model server run or performance sweep was performed in this worktree; no power limits were changed. The opt-in table is not a claim of speedup on untested cards.
Sur le site
Liens install, modèles, releases.