Issues / #1530

#1530 OCR character omission on a clear image with IQ3_S / GPU vision (v0.1.39 and v0.1.40.3)

open · @DamonFangCN · 3 Kommentare · Auf GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsSecurity

Beschreibung

## Summary

A clear synthetic confirmation code **`Z5HMXTSH`** is transcribed as **`Z5HMTSH`** (the `X` is omitted) with Qwen3.8-Flash-Next **IQ3_S / GPU vision**. This was observed on v0.1.39 and reproduced after upgrading to **v0.1.40.3** with unchanged model files and application configuration.

I am reporting a reproducible input/output case, **not claiming the cause is established to be a Strata bug**. I have not compared another engine, quantization, BF16 language model, MTP-off setting, or image-token budget. It could be a model/quantization limitation or involve the vision/inference path.

## Hardware and software

- Ubuntu **26.04.1 LTS**, x86-64
- **RTX 4090 24 GB**, **Ryzen 9 7950X**, **64 GB DDR5**
- NVIDIA driver **595.91.07**
- Local CUDA toolkit **13.4.1**; nvcc **13.4.59**
- Python **3.14.4**
- Strata **v0.1.40.3**, commit **`d5ea7133741e67743c0e886bb426c0ce8d69cf6c`**
- Locally built CUDA engine and GPU vision helper, compute capability **89**
- Build source hashes: language `3950c1f221aa6e1b`, vision `1fb2b3fb797ec402`
- llama.cpp pin **`3cf03257f219afbe7334045ff7c6a06ac68c627d`**

## Model and configuration

- Language model: **Qwen3.8-Flash-Next-GSQ-RCO, IQ3_S**, using the native `packs/iq3_s` data.
- GGUFs: `Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf` and `...-00002-of-00002.gguf`.
- GPU projector: **`mmproj-Qwen3.8-Flash-Next-BF16.gguf`**.
- Context **196,608**; KV **int8**; MTP **spec=4 / spec-min-p=0.5**; expert cache and prefill **auto**; VRAM reserve **700 MiB**.
- `vision.max_tokens=1024` (**image** token cap). The OCR request separately sets `max_tokens=512` (**generated output**, including thinking).
- Server bound to `127.0.0.1:8081`; API authentication enabled. Model ID: `qwen3.8-flash-next-iq3_s`.

<details>
<summary>Sanitized application configuration (paths replaced; actual secrets omitted)</summary>

```json
{
  "exe": "<STRATA_DIR>/engine/strata",
  "args": [
    "--pack",
    "<DATA_DIR>/packs/iq3_s",
    "--native",
    "<DATA_DIR>/models/IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf",
    "--ple-gguf",
    "<DATA_DIR>/models/IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf",
    "--expert-profile",
    "<STRATA_DIR>/data/expert-profile.bin",
    "--expert-cache",
    "auto",
    "--prefill",
    "auto",
    "--spec",
    "4",
    "--spec-min-p",
    "0.5",
    "--mtp",
    "<DATA_DIR>/mtp/rt",
    "--max-context",
    "196608",
    "--kv",
    "int8",
    "--vision",
    "--vram-reserve-mib",
    "700"
  ],
  "cwd": "<STRATA_DIR>",
  "tokenizer": "<DATA_DIR>/packs/iq3_s/tokenizer",
  "model_name": "qwen3.8-flash-next-iq3_s",
  "host": "127.0.0.1",
  "port": 8081,
  "open_browser": false,
  "gpu": 0,
  "gpus_asked": true,
  "lib_dirs": [
    "<STRATA_DIR>/.tools/cuda-13.4.1/bin",
    "<STRATA_DIR>/.tools/cuda-13.4.1/lib64"
  ],
  "vision": {
    "exe": "<STRATA_DIR>/engine/strata-vision",
    "mmproj": "<DATA_DIR>/models/vision/mmproj-Qwen3.8-Flash-Next-BF16.gguf",
    "model": "<DATA_DIR>/models/IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf",
    "gpu": true,
    "max_tokens": 1024
  },
  "api_key": "<REDACTED: authentication enabled>"
}
```

`<STRATA_DIR>` is the application directory; `<DATA_DIR>` is the Strata model-data directory on a dedicated model disk. The actual configuration and model files were unchanged across the upgrade.
</details>

## Input images

This is synthetic test data, not a private document. Both images visibly contain:

> 确认码:Z5HMXTSH

| Fixture | Dimensions | PNG bytes | SHA-256 |
|---|---|---:|---|
| `code-wide.png` | 2133 × 491 | 49,009 | `c6fc9d87257e5c96559021369eb2ceab108d9658e0e174c434770aa748c8b81f` |
| `code-tight.png` | 680 × 100 | 35,566 | `f113fd9d2a3dcf3eb2b15bafdaa29a67f45ef8dac96f9788607e9598fcb79e48` |

The tight image removes whitespace around the same line. Both are the **original archived PNGs**, not regenerated images. Byte-exact fixtures are supplied as compact, expandable Python snippets in the comments; running either snippet restores its PNG and checks its SHA-256. This text-based transfer avoids lossy screenshot conversion and does not require an external file host.

## Minimal reproduction

1. Restore the original PNG from the fixture comment below.
2. Use the configured local Strata server directly (no Open WebUI, reverse proxy, adapter, or tunnel).
3. Run this standard-library Python request. Set `STRATA_API_KEY` locally; do not paste a key into this issue.

```python
import base64, json, os, pathlib, urllib.request

image = pathlib.Path("code-wide.png")  # change to code-tight.png for that input
payload = {
    "model": "qwen3.8-flash-next-iq3_s",
    "max_tokens": 512,
    "stream": False,
    "messages": [{"role": "user", "content": [
        {"type": "text", "text": "读出图片中的确认码,按实际字形逐字符读取;只输出确认码,不猜测。"},
        {"type": "image_url", "image_url": {
            "url": "data:image/png;base64," + base64.b64encode(image.read_bytes()).decode()
        }}
    ]}]
}
# First batch: omit reasoning_effort, temperature, and seed, as in the old request.
# Auxiliary batch: add only payload["reasoning_effort"] = "none".
request = urllib.request.Request(
    "http://127.0.0.1:8081/v1/chat/completions",
    data=json.dumps(payload).encode(),
    headers={"Content-Type": "application/json",
             "Authorization": "Bearer " + os.environ["STRATA_API_KEY"]},
)
http = urllib.request.build_opener(urllib.request.ProxyHandler({}))
with http.open(request, timeout=120) as response:
    result = json.load(response)
print(json.dumps({"choices": result["choices"], "usage": result.get("usage")},
                 ensure_ascii=False, indent=2))
```

Expected final content: **`Z5HMXTSH`**, exactly, with no omitted character.

The prompt does not tell the model the correct answer or the expected number of characters. There is no system message, tool, or conversation history in these requests.

## Actual results on v0.1.40.3 (2026-10-08)

Three requests per image in each batch:

| Request settings | Wide image | Tight image |
|---|---|---|
| Default thinking; output cap 512; temperature/seed omitted | 3/3 `Z5HMTSH`; all `finish_reason=stop` | Once `Z5HMX TSH` (extra space); twice empty final content with `finish_reason=length`, 512 completion tokens |
| Same request, adding only `reasoning_effort="none"` | 3/3 `Z5HMTSH`; all `stop` | 3/3 `Z5HMTSH`; all `stop` |

Wide-image default-thinking runs: **2.391 / 0.606 / 0.512 seconds**, **73 / 70 / 72 completion tokens**. The missing `X` in these runs is not final-output truncation. First run: 1,080 prompt tokens, 0 cached tokens; subsequent runs report 1,075 cached prompt tokens.

Tight-image default-thinking runs: **1.045 / 4.098 / 4.135 seconds**. The first produced all characters but inserted whitespace (removing whitespace would recover the expected code). The other two hit the **test-specific 512-token ceiling**; they are not counted as explicit incorrect transcriptions. These are separate observations from the normally terminated missing-character case.

With thinking disabled, first wide/tight requests took **1.720 / 0.585 seconds**; repeats were faster with caches reused. **Repeated requests are not independent images or a population OCR-accuracy benchmark.** Sampling temperature and seed were not explicitly fixed; I am not asserting deterministic reproduction.

## Earlier observation on v0.1.39

The same wide PNG and same prompt sent directly to the local Strata API returned `Z5HMTSH` in **2.204 seconds**, `finish_reason=stop`, 73 completion tokens (1,080 prompt tokens; 0 cached).

An earlier public tight-crop request once returned the correct code, but used a different prompt: `逐字符读出图片里的确认码,只输出图片中实际的确认码,不要省略字符。` Historical sampling was not fixed. The current same-prompt wide/tight replay therefore does **not** establish a tight-crop accuracy regression, nor does one previous success make cropping a reliable fix. I have not switched back to v0.1.39 for a fresh strict A/B.

## What I would like help checking

1. Is this an expected limitation of this IQ3_S model, or a useful case to investigate in Strata's vision/inference path?
2. Would checking the image grid / minimum-vs-maximum image token behavior help? Related: #767 and #625, though this case uses **NVIDIA GPU vision with a 1024 cap**, not the CPU preset's 300 cap.
3. What is the most useful next comparison: same weights in llama.cpp, MTP off, another image-token budget, or another quantization?

No API keys, public service addresses, private documents, full configuration files with secrets, or unrelated service logs are included. Raw per-request metrics can be provided if useful.

Mehr auf der Site

Links zu Install, Modellen, Releases.