Issues / #1502
#1502 serve: a long prompt with images intermittently fails with "the image position upload failed" (#1242 family)
open · @PriceNing · 0 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux
Description
## Summary With `--vision` and MTP, a request whose prompt is long (37K+ tokens) and contains images intermittently fails during prompt processing: the engine reports `the image position upload failed`, and the request ends `done: 0 tokens (error, cancel=False)`. The same images on short prompts never fail. When it fires it tends to fire twice in a row for the same client request (two identical error lines at the same second); an identical retry later often passes. ## What happened 2026-10-08, engine 0.1.40.2, two client requests failed 70 s apart. Each request produced two identical engine error lines at the same second, then the error finish: ``` [strata] the engine reported an error: the image position upload failed [strata] done: 0 tokens in 0 s (0.0 tok/s) (error, cancel=False) [strata] the engine reported an error: the image position upload failed [strata] done: 0 tokens in 0 s (0.0 tok/s) (error, cancel=False) ``` No "reading the prompt" line appears for the failed requests: the failure is before the prompt read. ## Reproduction (intermittent, state-dependent) - Recipe: 37K-73K tokens of generated filler text + 1-2 small images + one short question. - Batch 1 (one engine process): 3 different prompts, 3/3 failed (two engine error lines each). - Batch 2, identical prompts, same process: 3/3 passed. After an engine restart: 3/3 passed. - Control: the same images on short prompts (incl. two concurrent image requests and one two-image request): 12/12 passed. Text-only requests fine throughout. No Xid/driver errors. ## Environment - engine 0.1.40.2 (v0.1.40.2 sources, CUDA 13.3, GCC 12.2), Linux - Qwen3.8-Flash-Next IQ3_XXS (512 experts/layer, top-k 10), `--max-context 262144 --kv int8 --kv-resident 20480 --spec 4 --spec-min-p 0.5 --mtp <draft> --vision --expert-cache auto --expert-cache-device1 3000 --conversation-cache-mib 12288 --conversation-cache-slots 6` - RTX 5060 Ti 16 GB (compute) + RTX 3060 12 GB (helper expert cache; the vision encoder runs on it) ## Notes The failure is the `cudaMemcpy` in `upload_mrope()` (generate.cpp:8815 in 0.1.40.2), after "positions built" / "device idle". A fixed-size H2D copy failing points at an earlier async CUDA fault in the same request. The engine resets the table to identity and keeps serving (the next text request is indeed clean), and identical retries often pass. This looks like the #1242 family (image-position tables vs captured decode graphs and MTP draft paths), but with a hard failure instead of a silent overwrite. Is this symptom covered by #1242's isolation, or is this a separate path? Happy to run this reproduction against a patch.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.