Pull requests / #1255

#1255 serve: /props total_slots is the engine's batch slots

closed · @midagedev · 0 comentarios · En GitHub

Server & APINVIDIA / CUDA

Descripción

`/props` always reports `total_slots: 1`. With `"parallel": 3` the engine runs three batch slots, and `/v1/status` reports `concurrency.serving: 3`. Clients that read `total_slots` think the server runs one request at a time. toktape refused a 3-session run for this reason.

The fix reports the engine's batch slots in `/props`, with the same expression as `/v1/status` (at least 1).

Evidence:
- RTX 3090 with `"parallel": 3`: the engine log says `--batch: 3 slot sessions`, but `/props` said `total_slots: 1`.
- New test `test_props_total_slots_follows_the_batch_slots`: fails on the current `main` (`1 != 3`) and passes with the fix. All 205 tests in `serve/test_server.py` pass.

`/slots` still lists one slot. That is a separate change.

This replaces #1004, which closed when `main` was force-pushed. The commit is the same, cherry-picked onto the new `main`.

I wrote this with an AI coding assistant and checked the change and the test runs myself.

En el sitio

Enlaces a install, modelos, releases.