Pull requests / #289

#289 Start: name the GPU and stop at once when the build has no code for it; STRATA_EMULATE_CC

closed · @sergqwer · 0 comentários · No GitHub

Setup & installNVIDIA / CUDA

Descrição

## Summary

- **A clear stop for a wrong build.** Before the expert arena loads, the engine names the card and its compute
  capability, and checks that the build has code for it (`cudaFuncGetAttributes` on a kernel). A binary built for
  other GPUs used to fail at its first kernel, after the whole arena had loaded. Now it stops in ~0.3 s, names the
  card and its `sm_`, and says to rebuild (setup does).
- **`STRATA_EMULATE_CC=75|80|86|89`**, for tests. With a PTX build for an older generation (e.g.
  `CMAKE_CUDA_ARCHITECTURES=86-virtual`, JIT-compiled on a newer card), the host's device queries answer as that
  card: compute capability for ggml's MMQ configs and the QSA kernel choice, shared memory per block (64 KB on
  Turing). The kernels are then the real card's, so their results are too, though their speed is not. This is how
  the fork tested RTX 20/30/40 code paths on an RTX 5090.
- **`STRATA_QSA_WARP=1|select|attn`**: the pre-sm_80 QSA kernels on any card, for A/B.

## Measured

On the fork, an sm_86-only build on the RTX 5090 stops in 0.3 s with the card's name. Emulated sm_75/86/89 builds
give first-token KL to the native sm_120a build of 0.0005 (sm_89), 0.0003 (sm_86) and 0.0034 (sm_75), all with the
same top-1. On main (IQ2_XS) the check prints the card and changes nothing else.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

No site

Links install, modelos, releases.