Pull requests / #660
#660 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
closed · @shefowl · 0 comentarios · En GitHub
Models & quantsDocumentationWindowsLinux
Descripción
On the CPU, strata-vision leaves mtmd's flash attention at `auto` (#288), which turns it on. Whether that is fast depends on how ggml was built. Its tiled CPU flash-attention kernel needs the head size (72 in this encoder) to be a multiple of the vector width: 8 floats with AVX2, 16 with AVX-512. A build from source on Zen 4 and newer, or on some Intel CPUs, gets AVX-512. There ggml falls back to a per-row kernel that is several times slower and accumulates V in FP16. The release builds are AVX2 (`STRATA_PORTABLE`), and on them flash attention is the faster option. This PR turns flash attention off by default only when the encoder runs on the CPU and its ggml has AVX-512 (`ggml_cpu_has_avx512()`). AVX2 builds, the release ones among them, keep `auto`. `--gpu` keeps `auto`, and an explicit `--flash-attn on|off|auto` still wins. The first version of this PR turned it off on every CPU. midhatn's measurement in the comments showed that this made the release build about 60% slower, which is how the cause turned up. ## Measured Ryzen 7 7700X (8 cores, AVX-512), Linux. `tools/vision` at the pinned llama.cpp, built natively and with `-DSTRATA_PORTABLE=ON`, with an F16 mmproj of Qwen3.8-Flash-Next. The test picture is the 1024x1024 one from the comments (1,024 image tokens), 8 threads, one encode per fresh process. Times are from the `OK` line and peak RSS from `VmHWM`. | Build | Flash attention | Encode | Peak RSS | Output vs FP32 attention | | --- | --- | ---: | ---: | ---: | | AVX2 (as released) | on (default before and after) | 8.3 / 9.9 s | 1,152 MiB | 1.2% | | AVX2 | off | 15.2 / 14.0 s | 2,194 MiB | 1.1% | | AVX-512 (from source) | on (default before) | 43.5 s | 1,152 MiB | 26% | | AVX-512 (from source) | off (default after) | 13.3 s | 2,195 MiB | reference | - The last column is the relative L2 distance of the embeddings from the AVX-512 run without flash attention. Two builds differ by about 1% anyway, so the tiled kernel loses nothing measurable. The fallback kernel is 26% off (cosine 0.965). - With this PR and no `--flash-attn`, the AVX2 build gave embeddings byte-identical to `on` (9.0 s), and the native build byte-identical to `off` (13.2 s). - llama.cpp master still has the same condition (`use_tiled &= (DV % f32_epr == 0)` in `ggml-cpu/ops.cpp`), so a newer pin would not change this. ## Changes - `tools/vision/strata_vision.cpp`: without `--flash-attn`, a CPU encoder whose ggml has AVX-512 runs with flash attention off. - `docs/DETAILS.md`: the cause and the table, under Images. Not measured: Windows, the BF16 mmproj, other CPUs. On an AVX-512 machine an encoder built without AVX-512 would be faster still (the first row of the table against the last). I left that build choice to you. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
En el sitio
Enlaces a install, modelos, releases.