Pull requests / #1207
#1207 fix(serve): keep cached images alive through request preparation
closed · @hulkbig · 0 comentarios · En GitHub
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
Descripción
Address the bounded image-cache lifetime gap reported in Niko1221/Strata#1072. Take the service FIFO before a reentrant Vision batch lock and keep both through encoding, marker expansion, validation, and the completed combined-file copy. Defer eviction until the outermost batch exits, then trim the successful raw-image cache back to 64 entries in finally. Repeated images reuse their encoding; downloads remain outside FIFO. Existing combined-file ownership is unchanged. AI-assisted implementation and independent AI review. CPU/Python model-free verification: 9 new regression methods and 15 existing vision tests pass; 10 external protocol probes pass, including a 65-image loopback HTTP 200 response and separate-process GENI file consumption. Independent verification passes 11 methods and 20 scheduled race repetitions. Focused 24-test suite rerun before publication: passed. Full server suite: 461 tests, 447 passed, 8 skipped, 6 errors from pre-existing missing regex imports, also observed on the base. No native CUDA/HIP inference, native C++ SVE loader, Windows, realistic embedding-volume stress, disk-full, or throughput validation. The longer FIFO hold is a correctness trade-off with unmeasured throughput cost. The separate encoder partial-write leak remains out of scope. Review branch only, based on 82f46a8c8f475f001ad76d92f58f4a4f8ffb0253.
En el sitio
Enlaces a install, modelos, releases.