Pull requests / #998
#998 serve: an engine that exits after an unread ERR line says why (#997 #890)
closed · @ischencheng · 0 コメント · GitHub で見る
本文
With `"parallel": N`, a batch window that fails prints why as an `ERR` line on stdout and the engine exits with code 1 (`batch_step` / the pipelined `pump` in generate.cpp, then `return 1` in the serve loop). At that point every request is reading its own slot's queue, so nobody reads that line. The server's note then falls back to the log's last `strata` line, which in #997 is the `expert tiers` stats line, and without a log it says "the usual cause is running out of RAM". `_pump` now keeps an `ERR` line that was the engine's last stdout line, and `death_note` uses it before the log-tail guess. With the test below the server window says: ``` [strata] the engine stopped unexpectedly (exit code 1). The engine exited (code 1) after it reported: verify batch: layer 34 never rang (graph finished) - please report it at github.com/Niko1221/Strata/issues with the log. The next request starts the engine again. ``` This doesn't fix whatever fails in the #997 / #890 windows. It only makes the next report name the error. If the engine dies some other way (a kernel wrapper's `std::exit` with only a stderr line, a signal), the note is the same as before. Test: `test_parallel`'s fake engine gets `--fail-window` (its first window over two slots prints an ERR line and exits 1). `test_a_failed_window_names_its_error` fails on main (the note blames RAM) and passes here. `serve/test_parallel.py`, `test_server.py`, `test_lifecycle.py`, `test_monitor.py`, `test_responses.py` and `test_vram.py` pass on macOS, Python 3.12. No engine change, so nothing was run on a GPU. Related to #997 and #890.
関連リンク
インストール・モデル・リリースへの站内リンク。