Issues / #1469
#1469 0.1.40.2: decode on a Tesla P40 (sm_61) is about half of 0.1.40's; bisected to 2e4ddf6, and restoring __restrict__ in pdl.hpp (STRATA_PDL_RESTRICT) gives back most of it
open · @paulhothersall · 2 コメント · GitHub で見る
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsLinux
本文
**On a Tesla P40 alone, v0.1.40.2 decodes at 17.7 tok/s against 36.4 on v0.1.40 (IQ3_XXS, 4K prompt, 9 vs 9 runs in 3 interleaved rounds, lower in every round) with prefill unchanged; a bisect over the 125 commits between the tags lands on 2e4ddf6 (the #904 port, "all opt-in"), and on that tag a one-line test change, `#define STRATA_PDL_RESTRICT __restrict__` (include/strata/kernels/pdl.hpp:41), brings decode back to 33.8, and together with `kPdlPrefetch = false` (line 47) to 36.7.** Limits: one rig, one quant, sm_61 only (the release notes' P100 is sm_60 and reads +3.7%); the two lines are a test, not a proposed patch, since pdl.hpp explains why those pointers lose `__restrict__` for PDL launches on sm_90+. Setup: Tesla P40 24 GB (cc 6.1), Ryzen 9 3900X, 64 GB DDR4-2400, PCIe 3.0 x8, Linux, CUDA 12.9, `-DCMAKE_CUDA_ARCHITECTURES="61;86"`, `--layer-split auto`, ctx 36,864, 4K prompt, 3 runs per cell, decode from the engine's per-run line. | build (P40 alone) | decode tok/s | share of the gap recovered | |---|---|---| | v0.1.40 | 36.4 (sd 1.6) | | | v0.1.40.2 | 17.7 (sd 1.0) | | | v0.1.40.2 + `STRATA_PDL_RESTRICT` = `__restrict__` | 33.8 (sd 1.7) | 86% | | v0.1.40.2 + `kPdlPrefetch = false` | 25.8 (sd 1.1) | 43% | | v0.1.40.2 + both | 36.7 (sd 3.4) | 100% | (Another interleaved set: 34.9 vs 18.3.) Prefill is the same (13.4-13.6 s TTFT in every arm). 256 tokens took 7.2 s on 0.1.40 and 14.3 s on 0.1.40.2 with about the same draft acceptance and GPU cache hit (93%), so each decode window takes about twice as long. On a P40 + RTX 3070 pair (3070 first, the P40 holds 5 layers) the A/B reads 36.7 vs 29.9 and 35.3 vs 30.7 (-13 to -18%); with the P40 first (16 layers) a single launch read 38 vs 27. Bisect (P40 alone, 3 runs per step): dfc8f74 36.3, 53623e5 36.8, bc089d5 36.7 good; 5e4de5c 17.6, 29897d7 18.4, **2e4ddf6 17.1** bad; 818ec1c 36.2, d6850f3 37.8 good. The commit message says every STRATA_DF_* switch is off by default; the slowdown is in the parts without a switch. My reading, not tested beyond the variants: on sm_61 `const __restrict__` is what lets the compiler use the read-only cache path, and the three kernels that lost it (`native_quantize_q8_1_kernel`, `bf16_f32_mmvf_multi_kernel`, the MMVQ multi-column kernel) run in every verify window. Ruled out: our extra sm_75 architecture (61;86 vs 61;75;86 build, same speed), `STRATA_UNBUFFERED_LOAD=1`, `STRATA_MMVQ_IL=0` (17.7 vs 17.8 on the P40 alone), `STRATA_NO_Q8K_AVX2=1` (17.8), reverting 7ef55e1 (#1118) and 48d2e6c (#1252) separately (no change), temperature and clocks (arms alternate within minutes, prefill identical), VRAM and cache layout. Checked in the source only: `include/strata/kernels/pdl.hpp` is unchanged in v0.1.40.3 (d5ea713), so lines 41 and 47 are still as in 0.1.40.2; I have not run 0.1.40.3. Not tested: other quants, other sm_61 cards (GTX 10 series), sm_60 / sm_70, a build with only 2e4ddf6 reverted, the effect on newer architectures.
関連リンク
インストール・モデル・リリースへの站内リンク。