Pull requests / #1255

#1255 serve: /props total_slots is the engine's batch slots

closed · @midagedev · 0 comments · View on GitHub

Server & APINVIDIA / CUDA

Description

`/props` always reports `total_slots: 1`. With `"parallel": 3` the engine runs three batch slots, and `/v1/status` reports `concurrency.serving: 3`. Clients that read `total_slots` think the server runs one request at a time. toktape refused a 3-session run for this reason.

The fix reports the engine's batch slots in `/props`, with the same expression as `/v1/status` (at least 1).

Evidence:
- RTX 3090 with `"parallel": 3`: the engine log says `--batch: 3 slot sessions`, but `/props` said `total_slots: 1`.
- New test `test_props_total_slots_follows_the_batch_slots`: fails on the current `main` (`1 != 3`) and passes with the fix. All 205 tests in `serve/test_server.py` pass.

`/slots` still lists one slot. That is a separate change.

This replaces #1004, which closed when `main` was force-pushed. The commit is the same, cherry-picked onto the new `main`.

I wrote this with an AI coding assistant and checked the change and the test runs myself.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.