Pull requests / #360
#360 qsa_prompt_attn: Volta (sm_70) crashes in the prompt path - compare the full compute capability
closed · @DingoOz · 0 comentarios · En GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsLinux
Descripción
## The bug With 0.1.31 built for a V100 (`-DCMAKE_CUDA_ARCHITECTURES=70 -DSTRATA_EXPERIMENTAL_SM60=ON`), the engine loads and answers a 22-token prompt, then dies on the first ~900-token prompt: ``` prefill copy_i32: unspecified launch failure ``` `qsa_prompt_attn_batch` reads only the compute capability's major version: `< 7` refuses the device, `< 8` means Turing. A V100 is 7.0, so it passes as Turing and launches the m16n8k8 kernel. In a sm_70 build `mma16816` compiles that kernel to `__trap()` (`__CUDA_ARCH__ < 750`), and the trap surfaces at the next sync as the launch failure above. The comment above the check already says an older card should keep the old kernel; only the check missed Volta. ## The fix Compare major*10+minor: below 75 returns false (the old kernel), below 80 is Turing. The variable is renamed from `cc_major` to `cc` because it now holds both parts. No change for sm_75 and newer, AMD, or Pascal (6.x was already refused). ## Tested Tesla V100-PCIE-16GB, CUDA 12.4, Linux, IQ3_S pack, `--kv int8 --spec 4`, greedy, 256 tokens out: | Prompt | Before | After (prompt / output tok/s) | | --- | --- | --- | | 22 tokens | OK | OK | | ~900 tokens | crash | 222 / 40 | | ~6.8K tokens | crash | 680 / 41.5 | | ~26.7K tokens | crash | 966 / 43 | These match 0.1.26 on the same card (221 / 680 / 982 tok/s prompt). The UD-Q4_K_XL import also runs with the fix (~20 tok/s output, with every expert in RAM).
En el sitio
Enlaces a install, modelos, releases.