Pull requests / #288

#288 strata-vision: --flash-attn switch, CPU runs without a CUDA context or warm-up; relative vision paths

closed · @sergqwer · 0 Kommentare · Auf GitHub

Server & APINVIDIA / CUDAModels & quants

Beschreibung

## Summary

Four small changes to the image encoder and its config:

- **`--flash-attn on|off|auto`** (mtmd's `flash_attn_type`). Flash attention keeps K and V in FP16. On the NVFP4 fork
  the vision tower's output was **11.2% off** with it on the GPU, and 0.1% with FP32 attention on the CPU. The
  default stays `auto`; the switch makes the choice possible.
- **CPU runs without a CUDA context.** A CUDA build of the encoder running on the CPU still opened a context on the
  GPU: 0.4-0.7 GB of VRAM, 150-260 expert slots less for the engine beside it. It now hides the GPU
  (`CUDA_VISIBLE_DEVICES=-1` before the runtime starts).
- **CPU threads and warm-up.** Without `--threads`, the CPU encoder uses one thread per core (mtmd's default is 4), and
  it skips the warm-up encode (~6 s of start at 1,024 image tokens).
- **Relative paths.** In `serve/server.py`, relative `exe`, `mmproj` and `model` paths in the vision config now
  resolve against the config's `cwd`, as the engine's paths already do.

## Measured

On the NVFP4 fork with an FP32 mmproj on the CPU: 2-6 s per picture, and no VRAM taken from the expert cache.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Mehr auf der Site

Links zu Install, Modellen, Releases.