Pull requests / #427
#427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)
closed · @homeofe · 0 commentaires · Sur GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Description
Follow-up to #411, #412 and #419. 0.1.33 fixed the shipped encoder in the release packaging step. This PR makes the repository itself safe by default, so that a release, CI job or hand build that forgets a flag can't ship an AVX-512 `strata-vision` again. It doesn't replace passing the flag in the release script; it's a second line of defence. There's no version bump; that's left for the release commit.
## Summary
| | before | after |
|---|---|---|
| `cmake -S tools/vision` with no flags | native (`GGML_NATIVE=ON`): the build machine's ISA, AVX-512 included | **portable**: AVX2 baseline |
| `-DSTRATA_PORTABLE=ON -DGGML_AVX512=ON` (or a hand-edited cache) | AVX-512 compiled in (10,522 `zmm` instructions) | AVX-512 forced OFF (0 `zmm`) |
| `setup.py` local encoder builds (CUDA `build_engine`, HIP `build_vision_cpu`) | native through the default | native through an explicit `-DSTRATA_PORTABLE=OFF` (**unchanged behaviour**) |
| `Dockerfile` encoder build | native through the default | native through an explicit `-DSTRATA_PORTABLE=OFF` (**unchanged behaviour**) |
| a build folder switched portable → native | kept `GGML_OPENMP=OFF` and the forced `GGML_AVX*=ON` in its cache | gets the same GGML cache entries as a fresh native configure |
| configure log | silent about the mode | `strata-vision: portable build …` or `strata-vision: native build for this CPU …` |
## Background: how the 0.1.32 encoder got AVX-512
The 0.1.32 `strata-windows-x64.zip` had a portable engine (`BUILD.json` says `"portable": true`) but a non-portable `strata-vision.exe`. #419's WinDbg trace shows an EVEX-encoded `vmovups zmm0, …` at `strata_vision+0x285545`, reached during model/backend init on the CPU path. It crashes with `0xC000001D` on every CPU without AVX-512: Intel 12th–14th gen consumer parts, Zen 2/3, 9th-gen Intel (#411: i9-9900KF; #412: Ryzen 9 5950X). A commenter on #419 confirmed that rebuilding with `-DSTRATA_PORTABLE=ON` fixes it.
The mechanism, in the pinned llama.cpp (`3cf03257`):
1. `tools/vision/CMakeLists.txt` had `option(STRATA_PORTABLE … OFF)`. A configure without the flag took the `else()` branch, `set(GGML_NATIVE ON CACHE BOOL "" FORCE)`.
2. On MSVC with `GGML_NATIVE`, `ggml/src/ggml-cpu/CMakeLists.txt:251-252` includes `ggml-cpu/cmake/FindSIMD.cmake`. That file compiles test programs on the **build** machine and does `set(GGML_AVX512 ON)` when `/arch:AVX512` works (`FindSIMD.cmake:95-100`).
3. `GGML_AVX512` then adds `/arch:AVX512` to ggml-cpu (`ggml-cpu/CMakeLists.txt:254-255`). On GCC/Clang the native branch is `-march=native` (`:307-308`), which has the same effect.
4. These flags are `PRIVATE` to the ggml-cpu target (`:675`), but ggml is linked statically into `strata-vision`, so the hot CPU kernels and init paths in the exe are AVX-512 code with no CPUID dispatch.
The engine doesn't have this problem because its own AVX-512 kernels (`src/kernels/cpu/iq_avx512.cpp`) are chosen at run time, and the release passes `-DSTRATA_PORTABLE=ON` for it. The encoder build relied on the same flag being passed, and in 0.1.32 it wasn't.
## Changes
### `tools/vision/CMakeLists.txt`
- **`STRATA_PORTABLE` defaults to `ON`.** A build with no flags is the AVX2 baseline: `GGML_NATIVE=OFF`, `GGML_AVX/AVX2/FMA/F16C/BMI2=ON`, `GGML_OPENMP=OFF`, and on MSVC the static runtime (`CMAKE_MSVC_RUNTIME_LIBRARY=MultiThreaded`), exactly as with `-DSTRATA_PORTABLE=ON` before.
- **Portable now forces everything above AVX2 OFF:** `GGML_AVX512`, `GGML_AVX512_VBMI`, `GGML_AVX512_VNNI`, `GGML_AVX512_BF16`, `GGML_AVX_VNNI`, `GGML_AMX_TILE`, `GGML_AMX_INT8`, `GGML_AMX_BF16`. All eight are real options in the pinned ggml (`ggml/CMakeLists.txt:156-170`; the AMX ones only exist outside MSVC, where the cache entries are unused and harmless). `GGML_AVX_VNNI` is included because a 9900K, a Zen 2 or a Zen 3 has AVX2 but no AVX-VNNI. Before this, only ggml's defaults kept these OFF, so `-DGGML_AVX512=ON` (or an edited cache) got through even with `-DSTRATA_PORTABLE=ON`.
- **Native (`STRATA_PORTABLE=OFF`) now restores a fresh native configure's state.** The portable branch FORCE-writes `GGML_OPENMP=OFF` and `GGML_AVX*=ON` into the cache, and the old `else()` branch never undid that. A folder configured portable and then reused by `setup.py` (which passes `OFF`) therefore stayed without OpenMP. On MSVC it also kept `GGML_BMI2=ON`, which adds `__BMI2__` next to FindSIMD's own detection. The `else()` branch now sets `GGML_OPENMP=ON` and `GGML_AVX/AVX2/FMA/F16C/BMI2=OFF`, ggml's own defaults for a native build. I verified below that the GGML cache entries match a fresh native configure exactly.
- **The configure log names the mode** (`message(STATUS …)`), so the one case this PR can't fix is visible in every build log (see *Limitations*).
- The header documents the default, the opt-out and the cache caveat.
### `setup.py`
Both places that compile the encoder for the PC they run on pass `-DSTRATA_PORTABLE=OFF`, so local installs keep native codegen exactly as before:
- `build_vision_cpu` (`setup.py:1196`): the CPU encoder beside the HIP engine (#304).
- `build_engine` (`setup.py:1534-1536`): the encoder beside the CUDA engine, for `--vision gpu` and `--vision cpu`.
### `Dockerfile`
The encoder build (`Dockerfile:84-86`) passes `-DSTRATA_PORTABLE=OFF`. The image is built on the PC it runs on (README: `docker build … -t strata` then `docker run … strata`), and the engine in the same image is native too (the root `STRATA_PORTABLE` default is still OFF), so the image behaves the same as before. The comment above the build step now says that the image's code is native to the CPU that builds it.
### Tests (`tools/test_setup_choices.py`)
- `HipVision.test_the_cpu_encoder_is_built_once` also asserts `-DSTRATA_PORTABLE=OFF` for `build_vision_cpu`.
- New `CudaVision.test_the_encoder_is_built_native` covers `build_engine`. It runs a temp `ROOT` with an engine already built (`BUILD.json` `source=local`, matching `src`), so only the encoder is compiled, and mocks `cmake_build`, `source_hash`, `install_build_tools` and `source_version`. For `vision` in (`gpu`, `cpu`) it asserts:
- the only target is `strata-vision`
- `-DSTRATA_PORTABLE=OFF` is passed
- `-DSTRATA_VISION_CUDA` is `ON` for gpu and `OFF` for cpu
**It fails on `main`'s `setup.py` (2 failures) and passes here**, so it guards against the flag being dropped in future.
## Verification
### Test machine
| | |
|---|---|
| CPU | AMD Ryzen Threadripper 3960X (Zen 2), 24 cores / 48 threads, up to 4.57 GHz, 128 MiB L3 |
| CPU features | `sse4_2 avx2 fma f16c bmi2`; **no AVX-512, no AVX-VNNI**. This is the class of CPU the 0.1.32 encoder crashed on (#411 i9-9900KF, #412 Ryzen 9 5950X, #419 i7-13700KF) |
| RAM | 128 GB DDR4-2133, 8 × 16 GB (125 GiB usable), 8 GiB swap |
| GPU | NVIDIA GeForce RTX 2080 Ti, 11 GiB (sm_75), driver 595.84, PCIe 3.0 x16 |
| OS | Ubuntu 24.04.4 LTS, kernel 6.8.0 |
| Toolchain | GCC 13.3.0, CMake 3.28.3, Ninja 1.11.1, GNU binutils/objdump 2.42, Python 3.12.3, CUDA 12.0 (nvcc 12.0.140) |
**Load during the tests.** This PC runs Strata day to day, so the tests ran next to a live model server and not on an idle machine:
- **Running model:** Qwen3.8-Flash-Next IQ3_S on a local 0.1.31 engine, 128K context, `--kv int8`, `--spec 4`.
- **Engine resources:** about 51 GB resident (≈47 GiB of experts in RAM). The GPU was at 10.5 of 11 GiB, and system RAM at about 58 of 125 GiB used.
- **Concurrency:** the engine kept answering requests while the test builds ran with `-j 40`.
The builds in the table are CPU-only, so they didn't use the GPU. Load affects build time only, not the generated code or the `zmm` counts.
### Builds
Each case is a CPU-only `strata-vision` build:
```
cmake -G Ninja -S tools/vision -B <dir> -DCMAKE_BUILD_TYPE=Release \
-DLLAMA_DIR=third_party/llama.cpp -DSTRATA_VISION_CUDA=OFF [extra flags]
cmake --build <dir> --target strata-vision
objdump -d --no-show-raw-insn <dir>/bin/strata-vision | grep -c zmm
```
| # | CMakeLists | extra flags | ggml-cpu flags (`Adding CPU backend variant ggml-cpu: …`) | `zmm` |
|---|---|---|---|---|
| 1 | `main` | none | `-march=native` (the 0.1.32 mode; on an AVX-512 machine with MSVC this is `/arch:AVX512`) | – (configure only) |
| 2 | `main` | `-DSTRATA_PORTABLE=ON -DGGML_AVX512=ON` | `-msse4.2 -mf16c -mfma -mbmi2 -mavx -mavx2 -mavx512f -mavx512cd -mavx512vl -mavx512dq -mavx512bw` | **10,522** |
| 3 | this PR | none | `-msse4.2 -mf16c -mfma -mbmi2 -mavx -mavx2` | **0** |
| 4 | this PR | same folder as 3, `-DGGML_AVX512=ON -DGGML_AVX512_VNNI=ON -DGGML_AVX_VNNI=ON` | same as 3; all three cache entries end up `OFF` | **0** |
| 5 | this PR | same folder, `-DSTRATA_PORTABLE=OFF` | `-march=native`; log says `native build for this CPU` | – |
| 6 | this PR | fresh folder, `-DSTRATA_PORTABLE=OFF` | `-march=native` | – |
| 7 | this PR | folder 5 switched back to `-DSTRATA_PORTABLE=ON` and rebuilt | same as 3 | **0** |
- In case 2, the `zmm` code is in the hot paths an AVX2-only CPU runs during real work: `ggml_gemm_q2_K_8x8_q8_K` (2,303 lines), `ggml_gemm_q4_K_8x8_q8_K` (1,061), the `gemm_q4_b32_8x8_q8_0_lut_avx` templates (719 each) and `ggml_compute_forward_flash_attn_ext` (410).
- Cases 5 and 6 have **identical** GGML cache entries: `GGML_NATIVE=ON`, `GGML_OPENMP=ON`, `GGML_AVX/AVX2/FMA/F16C/BMI2/AVX512=OFF`.
- `ldd` on the binary shows only `libstdc++`, `libm`, `libgcc_s` and `libc`, so ggml is linked statically and the `zmm` count covers all of it. The binary prints its usage and exits 2 when run with no arguments.
Unit tests:
```
python3 -m unittest tools.test_setup_amd tools.test_setup_choices tools.test_setup_draft_vocab \
tools.test_setup_lowram tools.test_setup_pins tools.test_setup_prompts tools.test_setup_risk \
tools.test_setup_rope tools.test_setup_unsloth # all OK (choices: 16 tests incl. CudaVision)
```
`tools.test_setup_golden` fails 46 times on Linux both on `main` and here. That's an unrelated normalizer bug, fixed in #428.
I couldn't test on this Linux box: MSVC/Windows, CUDA encoder builds, and an AVX-512 build machine. The Windows behaviour follows from the ggml code cited above: with `GGML_NATIVE=OFF`, FindSIMD isn't included and only the cache options decide, and those are now forced.
### Also: the CUDA encoder on this PC
After the PR checks, I updated this PC's own server to 0.1.33 with the image encoder built the way `setup.py` does it, with this PR's flag:
```
cmake -S tools/vision -B build-vision -G Ninja -DCMAKE_BUILD_TYPE=Release -DLLAMA_DIR=third_party/llama.cpp \
-DSTRATA_VISION_CUDA=ON -DSTRATA_PORTABLE=OFF -DCMAKE_CUDA_ARCHITECTURES=75 -DCMAKE_CUDA_COMPILER=/usr/bin/nvcc
```
That build used `main`'s `tools/vision/CMakeLists.txt` with the flag passed explicitly, which is what this PR's `setup.py` passes.
- **Build:** CUDA 12.0 for sm_75. The cache shows `GGML_NATIVE=ON`, `GGML_CUDA=ON`, `STRATA_PORTABLE=OFF`.
- **Config:** the encoder runs on the 2080 Ti beside the engine (`--vision --vram-reserve-mib 700`). It takes 1.26 GiB of VRAM, and the engine's GPU expert cache goes from 1,148 to 509 experts.
- **Results:**
- A test image (a red square and a blue circle) was described correctly ("a red square on the left and a blue circle on the right") in 2.5 s end to end.
- Text decoding stayed at 35.0–35.3 tok/s, against 33.8–36.8 tok/s before.
## Impact by user group
- **Ready-made Windows zip users:** no change from this PR by itself. The next release's encoder is portable even if the script forgets the flag, provided the build folder is fresh (see *Limitations*).
- **Local CUDA installs (`setup.py`, Windows and Linux):** still native. `VISION_SOURCES = ("tools/vision",)` hashes the folder, so `vision_src` changes and an installed encoder is rebuilt **once** at the next setup run or update, like any other change under `tools/vision`. The rebuilt binary is the same native build as before.
- **Local HIP installs with the CPU encoder Sur le site
Liens install, modèles, releases.