Issues / #892
#892 Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works
open · @Thanh-Huy1104 · 2 commentaires · Sur GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsLinux
Description
## What happens On Linux, setup finds no ready-made engine and compiles one with whatever toolkit is installed. With CUDA 13.2 on an RTX 5070 Ti, the engine loads and serves normally, but every answer is wrong: - "What is 2+2?" -> `4+4=8.8.` - "Translate 'good morning' into German." -> `Good morning` - Most other prompts end after 1-2 tokens or loop (`primary colors are primary colors are ...`). There are no errors in the log. The benchmark script's warm-up fails with `no text received`. ## Environment - RTX 5070 Ti 16 GB (sm_120), driver 580.178.04, Ubuntu 24.04 (kernel 6.8.0-142), Core Ultra 7 265K (AVX2, no AVX-512), 62 GB RAM (KVM VM, GPU passthrough) - Strata `6f32ec0` (engine 0.1.39), `./setup.sh --yes --family qwen --model IQ2_XS --no-start` - Setup compiled with `/usr/local/cuda-13.2/bin/nvcc` (13.2.51) - Model GGUF SHA-256s match the published LFS hashes (same as the RTX 5090 community report) ## Isolation | Run | Result | | --- | --- | | Strata engine, CUDA 13.2 (any `--pcie-frac`, `--expert-cache 0`, 4K or 64K context) | garbage | | llama.cpp `3cf0325` (Strata's pinned commit), CUDA 13.2, GPU offload, same GGUF | garbage (0/12 on a small task suite) | | llama.cpp, same build, `-ngl 0 --device none` (CPU only) | correct | | llama.cpp `test-backend-ops -b CUDA0`, CUDA 13.2 | `MUL_MAT` 44 FAIL, `MUL_MAT_ID` 22 FAIL: **only** `iq1_s`, `iq2_s`, `iq3_s` (ERR 0.4-0.66); every other type OK | | **Strata engine rebuilt with CUDA 13.0.88** (pip `nvidia-cuda-nvcc==13.0.88`), nothing else changed | **correct**: 12/12 on the task suite | This matches the known nvcc 13.2 / sm_120 miscompile of the byte reads in the IQ1_S/IQ2_S/IQ3_S kernels (ggml-org/llama.cpp#21255, #28581; the fix PR ggml-org/llama.cpp#28784 was closed unmerged). Strata's own `src/kernels/cuda/iq_kernels.cu` uses the same pattern (`const uint8_t* qs = (const uint8_t*) &qs_packed;` then `qs[l]` into `iq2s_grid` / `iq3s_grid` / `iq1s_grid_gpu`, e.g. lines 156, 215, 242, 658, 719, 753). I have not tested which of Strata's kernels and ggml's kernels are affected separately. A PTX-only build (`120-virtual`) with 13.2 is not a workaround on driver 580: `the provided PTX was compiled with an unsupported toolchain`. ## Results with the CUDA 13.0 engine (for reference) `bench/results/2026-09-30-community-rtx-5090/benchmark.py`, 3 runs, 256 output tokens, 0 reused: | Prompt tokens | Prompt tok/s | Decode tok/s | | ---: | ---: | ---: | | 4,096 | 2,706 | 128 | | 32,768 | 3,280 | 130 | | 60,000 | 3,196 | 125 | ## Suggestions 1. Setup: when the toolkit is nvcc 13.2 and a card is sm_120, warn or stop. Alternatively, compile with pip's `nvidia-cuda-nvcc==13.0.88` (plus `nvidia-cuda-cccl`, `nvidia-nvvm`, `nvidia-cuda-crt`, and the runtime/cuBLAS wheels setup already pins). That works without sudo. 2. Kernels: replace the `uint8_t*` byte indexing with `__byte_perm` / shifts in the IQ*_S paths (what ggml-org/llama.cpp#28784 did). 3. A one-line self-check after the first start (for example, "capital of France" -> contains "Paris") would catch a broken build before users rely on it.
Sur le site
Liens install, modèles, releases.