Issues / #971

#971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)

closed · @talisp · 0 comentários · No GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Descrição

`docs/UNSLOTH_Q4.md` says UD-IQ4_XS with images is "not yet run with images". Here it runs, and it is in daily agent use (Hermes Agent on Windows, over the OpenAI API, driving Unreal Engine work). This report covers what we did to make it run in a VM with a vGPU on an older driver, so others (people and LLMs) can repeat it.

## Hardware

| | |
|---|---|
| GPU | NVIDIA **L40S-48Q vGPU** (VMware ESXi): 48 GB reported, **~42.2 GiB actually allocatable** (~5.5 GiB is vGPU/hypervisor overhead) |
| Driver | **550.127.05** (exposes CUDA 12.4) |
| CPU | Intel Xeon Platinum 8580, 32 vCPU, AVX-512 |
| RAM | 251 GB |
| OS | Ubuntu 22.04.5 LTS (kernel 6.8), a VM |

## Running a CUDA 13 project on a CUDA 12 driver

In a vGPU guest, the NVIDIA driver is coupled to the host's vGPU Manager, so **it cannot be upgraded from inside the VM**. CUDA 13 needs driver 580+, so here CUDA 12.4 is the ceiling.

- **0.1.38:** setup refused drivers below 580 (`MIN_DRIVER = 580`). We installed the **CUDA 12.4 toolkit in a user folder** (no system install), built the engine and `strata-vision` from source (`--build`, `sm_89`), and ran setup through a small wrapper. The wrapper put CUDA 12.4 first on `PATH` / `LD_LIBRARY_PATH` and lowered `MIN_DRIVER` to 550 **in memory only**, without editing the repo. Everything ran on driver 550 via CUDA minor-version compatibility.
- **0.1.39:** this is official now, so the wrapper only sets the toolkit path. `--cuda 12` accepts driver 525+ on Linux:

```
setup.py --setup --family unsloth --model UD-IQ4_XS --context 262144 --kv int8 --vision gpu --cuda 12 --build --no-start --yes
```

The engine went to `engine-cuda12/`, so the earlier IQ3_S install stayed intact for rollback. The encoder is the same `mmproj-Qwen3.8-Flash-Next-BF16.gguf` (SHA-256 checked). Engine log: all 55 GiB of experts in RAM; expert cache 14,672 slots (33.1 GiB of VRAM); 43.4 GB of VRAM in use with images on.

## VM / vGPU limitations worth knowing

- **Plan VRAM against ~42.2 GiB, not 48 GB.** Tools that size themselves as a fraction of "total" overshoot.
- **The driver cannot be upgraded from the guest**: CUDA 13 builds and the ready-made engine are out; the path is CUDA 12 + `--build`.
- **The GPU may be shared** with other services in the same VM: only one GPU workload at a time fits.
- **Some telemetry is missing on vGPU** (power, temperature).

## Checks (all passed)

- `/health` reports `images: true` at 262144 context.
- **A synthetic image** (3 shapes, 3 colours): correct count, shapes and order.
- **A Blender viewport render** (a mannequin on a throne): correct count, where the hands are, where the feet are.
- **In the agent:** yes/no questions about Unreal viewport captures agreed with pixel measurements. On an all-black capture, the model said it could not see the object instead of inventing it.
- **Also on the same build:** tool calls, streaming, and finding a fact at 89.5K tokens of context.

## Speed: a small, fair trade

Same machine, images on in both:

| | IQ3_S (0.1.38) | UD-IQ4_XS (0.1.39) |
|---|---|---|
| Prompt reading at ~90K context | ~3,950 tok/s | ~3,300 tok/s |
| Decode at ~90K context | ~93 tok/s | ~84 tok/s |
| Request with one image (incl. answer) | — | 3.6–4.8 s |

About 10% slower decode and ~17% slower prompt reading. In our use, the better answers and visual judgements of the ~4-bit model are worth it. This is a usage impression, not a controlled A/B: a small blind comparison of text and reasoning did not show a measurable difference yet. Some small "filling-in" was seen (4 window panes where there were 2, shading read as body features), so we keep numeric checks as the primary gate and vision as a second opinion.

## Tips

- **Keep your client's model name:** add `"aliases": ["<old model name>"]` to the new `strata-*.json`, so a client configured with that name needs no change.
- **Carry your edits over:** setup writes a new `strata-*.json` per model. Copy your own edits (`api_key`, `host`, `sampling`, `reasoning_budget_tokens`) from the old one.
- **The `<think>` fix:** see #804 / #970 for the fix for tool calls written inside the thinking (seen with this model family in agent use).

Next, I will test even larger quantizations on this machine and report back.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

No site

Links install, modelos, releases.