Issues / #892

#892 Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works

open · @Thanh-Huy1104 · 2 commentaires · Sur GitHub

BenchmarksSetup & installNVIDIA / CUDAModels & quantsLinux

Description

## What happens

On Linux, setup finds no ready-made engine and compiles one with whatever toolkit is installed. With CUDA 13.2 on an RTX 5070 Ti, the engine loads and serves normally, but every answer is wrong:

- "What is 2+2?" -> `4+4=8.8.`
- "Translate 'good morning' into German." -> `Good morning`
- Most other prompts end after 1-2 tokens or loop (`primary colors are primary colors are ...`).

There are no errors in the log. The benchmark script's warm-up fails with `no text received`.

## Environment

- RTX 5070 Ti 16 GB (sm_120), driver 580.178.04, Ubuntu 24.04 (kernel 6.8.0-142), Core Ultra 7 265K (AVX2, no AVX-512), 62 GB RAM (KVM VM, GPU passthrough)
- Strata `6f32ec0` (engine 0.1.39), `./setup.sh --yes --family qwen --model IQ2_XS --no-start`
- Setup compiled with `/usr/local/cuda-13.2/bin/nvcc` (13.2.51)
- Model GGUF SHA-256s match the published LFS hashes (same as the RTX 5090 community report)

## Isolation

| Run | Result |
| --- | --- |
| Strata engine, CUDA 13.2 (any `--pcie-frac`, `--expert-cache 0`, 4K or 64K context) | garbage |
| llama.cpp `3cf0325` (Strata's pinned commit), CUDA 13.2, GPU offload, same GGUF | garbage (0/12 on a small task suite) |
| llama.cpp, same build, `-ngl 0 --device none` (CPU only) | correct |
| llama.cpp `test-backend-ops -b CUDA0`, CUDA 13.2 | `MUL_MAT` 44 FAIL, `MUL_MAT_ID` 22 FAIL: **only** `iq1_s`, `iq2_s`, `iq3_s` (ERR 0.4-0.66); every other type OK |
| **Strata engine rebuilt with CUDA 13.0.88** (pip `nvidia-cuda-nvcc==13.0.88`), nothing else changed | **correct**: 12/12 on the task suite |

This matches the known nvcc 13.2 / sm_120 miscompile of the byte reads in the IQ1_S/IQ2_S/IQ3_S kernels (ggml-org/llama.cpp#21255, #28581; the fix PR ggml-org/llama.cpp#28784 was closed unmerged). Strata's own `src/kernels/cuda/iq_kernels.cu` uses the same pattern (`const uint8_t* qs = (const uint8_t*) &qs_packed;` then `qs[l]` into `iq2s_grid` / `iq3s_grid` / `iq1s_grid_gpu`, e.g. lines 156, 215, 242, 658, 719, 753). I have not tested which of Strata's kernels and ggml's kernels are affected separately.

A PTX-only build (`120-virtual`) with 13.2 is not a workaround on driver 580: `the provided PTX was compiled with an unsupported toolchain`.

## Results with the CUDA 13.0 engine (for reference)

`bench/results/2026-09-30-community-rtx-5090/benchmark.py`, 3 runs, 256 output tokens, 0 reused:

| Prompt tokens | Prompt tok/s | Decode tok/s |
| ---: | ---: | ---: |
| 4,096 | 2,706 | 128 |
| 32,768 | 3,280 | 130 |
| 60,000 | 3,196 | 125 |

## Suggestions

1. Setup: when the toolkit is nvcc 13.2 and a card is sm_120, warn or stop. Alternatively, compile with pip's `nvidia-cuda-nvcc==13.0.88` (plus `nvidia-cuda-cccl`, `nvidia-nvvm`, `nvidia-cuda-crt`, and the runtime/cuBLAS wheels setup already pins). That works without sudo.
2. Kernels: replace the `uint8_t*` byte indexing with `__byte_perm` / shifts in the IQ*_S paths (what ggml-org/llama.cpp#28784 did).
3. A one-line self-check after the first start (for example, "capital of France" -> contains "Paris") would catch a broken build before users rely on it.

Sur le site

Liens install, modèles, releases.