Pull requests / #360

#360 qsa_prompt_attn: Volta (sm_70) crashes in the prompt path - compare the full compute capability

closed · @DingoOz · 0 comments · View on GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsLinux

Description

## The bug

With 0.1.31 built for a V100 (`-DCMAKE_CUDA_ARCHITECTURES=70 -DSTRATA_EXPERIMENTAL_SM60=ON`), the engine loads and
answers a 22-token prompt, then dies on the first ~900-token prompt:

```
prefill copy_i32: unspecified launch failure
```

`qsa_prompt_attn_batch` reads only the compute capability's major version: `< 7` refuses the device, `< 8` means
Turing. A V100 is 7.0, so it passes as Turing and launches the m16n8k8 kernel. In a sm_70 build `mma16816` compiles that
kernel to `__trap()` (`__CUDA_ARCH__ < 750`), and the trap surfaces at the next sync as the launch failure above. The
comment above the check already says an older card should keep the old kernel; only the check missed Volta.

## The fix

Compare major*10+minor: below 75 returns false (the old kernel), below 80 is Turing. The variable is renamed from
`cc_major` to `cc` because it now holds both parts. No change for sm_75 and newer, AMD, or Pascal (6.x was already
refused).

## Tested

Tesla V100-PCIE-16GB, CUDA 12.4, Linux, IQ3_S pack, `--kv int8 --spec 4`, greedy, 256 tokens out:

| Prompt | Before | After (prompt / output tok/s) |
| --- | --- | --- |
| 22 tokens | OK | OK |
| ~900 tokens | crash | 222 / 40 |
| ~6.8K tokens | crash | 680 / 41.5 |
| ~26.7K tokens | crash | 966 / 43 |

These match 0.1.26 on the same card (221 / 680 / 982 tok/s prompt). The UD-Q4_K_XL import also runs with the fix
(~20 tok/s output, with every expert in RAM).

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.