Pull requests / #1004

#1004 serve: /props total_slots is the engine's batch slots

closed · @midagedev · 0 评论 · 在 GitHub 查看

Server & APINVIDIA / CUDA

描述

`/props` always reports `total_slots: 1`. With `"parallel": 3` the engine runs three batch slots, and `/v1/status` reports `concurrency.serving: 3`. Clients that read `total_slots` think the server runs one request at a time. toktape refused a 3-session run for this reason.

The fix reports the engine's batch slots in `/props`, with the same expression as `/v1/status` (at least 1).

Evidence:
- RTX 3090 with `"parallel": 3`: the engine log says `--batch: 3 slot sessions`, but `/props` said `total_slots: 1`.
- New test `test_props_total_slots_follows_the_batch_slots`: fails before the fix (`1 != 3`) and passes after. All 140 tests in `serve/test_server.py` pass.

`/slots` still lists one slot. Open PR #562 changes that block, so I left it alone.

I wrote this with an AI coding assistant and checked the change and the test runs myself.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。