Pull requests / #1004

#1004 serve: /props total_slots is the engine's batch slots

closed · @midagedev · 0 コメント · GitHub で見る

Server & APINVIDIA / CUDA

本文

`/props` always reports `total_slots: 1`. With `"parallel": 3` the engine runs three batch slots, and `/v1/status` reports `concurrency.serving: 3`. Clients that read `total_slots` think the server runs one request at a time. toktape refused a 3-session run for this reason.

The fix reports the engine's batch slots in `/props`, with the same expression as `/v1/status` (at least 1).

Evidence:
- RTX 3090 with `"parallel": 3`: the engine log says `--batch: 3 slot sessions`, but `/props` said `total_slots: 1`.
- New test `test_props_total_slots_follows_the_batch_slots`: fails before the fix (`1 != 3`) and passes after. All 140 tests in `serve/test_server.py` pass.

`/slots` still lists one slot. Open PR #562 changes that block, so I left it alone.

I wrote this with an AI coding assistant and checked the change and the test runs myself.

関連リンク

インストール・モデル・リリースへの站内リンク。