Pull requests / #562
#562 serve: /slots reports n_prompt_tokens, so a llama.cpp-style context m…
closed · @Zerschranzer · 0 Kommentare · Auf GitHub
Server & APIMulti-GPUAMD / HIPModels & quantsDocumentation
Beschreibung
## Why
Clients written against llama.cpp's endpoints read a context meter from `/slots`: they divide `n_prompt_tokens` by
`n_ctx`. Strata's slot carried only `id`, `n_ctx` and `is_processing`, so such a meter sits at **0 %** against
Strata however full the context is - and the reaction a user reaches for ("raise the context") changes nothing.
The number is already there, only not under that name. Same second, IQ3_S, 2x RX 9060 XT (gfx1200, HIP), engine
0.1.37:
GET /slots -> [{"id": 0, "n_ctx": 262144, "is_processing": false}]
GET /status -> {"busy": false, ..., "prompt_tokens": 120711, "generated": 331}
I hit it with a chat front-end of my own that draws a context ring: it stayed at 0 % through a 120K-token
conversation, while `/status` and `/metrics` had the right number the whole time.
## What it changes
- `GET /slots` also returns `n_prompt_tokens`: the running request's prompt size, kept after it ends (the
conversation cache carries it on). Read under `status_lock`, so it is the same snapshot `/status` serves.
Nothing else in the response changes, and the Monitor's own `/metrics` is untouched.
- `serve/test_server.py`: `test_slot_reports_the_context_in_use` - a set `prompt_tokens` shows up in the slot, and
after clearing it the slot reports 0. The existing exact-slot assertion in
`test_build_model_path_and_slot_status` carries the new key.
- `docs/DETAILS.md`: the endpoint table and the paragraph say what `/slots` reports.
## Semantics, on purpose
- Generated tokens are not counted, exactly as llama.cpp's `n_prompt_tokens` does not count them; the next
request's prompt carries them through the conversation cache, so a meter catches up on the next round.
- The value is set when the request starts, so it jumps to the full prompt size at prefill instead of growing
token by token.
- A dead engine still answers `[]`, so a client hides the meter rather than showing 0 %.
- `n_prompt_tokens_processed` and `n_prompt_tokens_cache` (also llama.cpp names) are deliberately left out until
someone needs them; the engine's `reused` would map onto the latter.
## Tests
python -m unittest serve.test_server serve.test_monitor
Ran 122 tests in 53.453s - OK
against the mock engine, no GPU and no pack. Live on the same build:
GET /slots -> [{"id": 0, "n_ctx": 262144, "is_processing": false, "n_prompt_tokens": 120711}]
Mehr auf der Site
Links zu Install, Modellen, Releases.