Pull requests / #562

#562 serve: /slots reports n_prompt_tokens, so a llama.cpp-style context m…

closed · @Zerschranzer · 0 comentários · No GitHub

Server & APIMulti-GPUAMD / HIPModels & quantsDocumentation

Descrição

## Why

Clients written against llama.cpp's endpoints read a context meter from `/slots`: they divide `n_prompt_tokens` by
`n_ctx`. Strata's slot carried only `id`, `n_ctx` and `is_processing`, so such a meter sits at **0 %** against
Strata however full the context is - and the reaction a user reaches for ("raise the context") changes nothing.

The number is already there, only not under that name. Same second, IQ3_S, 2x RX 9060 XT (gfx1200, HIP), engine
0.1.37:

    GET /slots   -> [{"id": 0, "n_ctx": 262144, "is_processing": false}]
    GET /status  -> {"busy": false, ..., "prompt_tokens": 120711, "generated": 331}

I hit it with a chat front-end of my own that draws a context ring: it stayed at 0 % through a 120K-token
conversation, while `/status` and `/metrics` had the right number the whole time.

## What it changes

- `GET /slots` also returns `n_prompt_tokens`: the running request's prompt size, kept after it ends (the
  conversation cache carries it on). Read under `status_lock`, so it is the same snapshot `/status` serves.
  Nothing else in the response changes, and the Monitor's own `/metrics` is untouched.
- `serve/test_server.py`: `test_slot_reports_the_context_in_use` - a set `prompt_tokens` shows up in the slot, and
  after clearing it the slot reports 0. The existing exact-slot assertion in
  `test_build_model_path_and_slot_status` carries the new key.
- `docs/DETAILS.md`: the endpoint table and the paragraph say what `/slots` reports.

## Semantics, on purpose

- Generated tokens are not counted, exactly as llama.cpp's `n_prompt_tokens` does not count them; the next
  request's prompt carries them through the conversation cache, so a meter catches up on the next round.
- The value is set when the request starts, so it jumps to the full prompt size at prefill instead of growing
  token by token.
- A dead engine still answers `[]`, so a client hides the meter rather than showing 0 %.
- `n_prompt_tokens_processed` and `n_prompt_tokens_cache` (also llama.cpp names) are deliberately left out until
  someone needs them; the engine's `reused` would map onto the latter.

## Tests

    python -m unittest serve.test_server serve.test_monitor
    Ran 122 tests in 53.453s - OK

against the mock engine, no GPU and no pack. Live on the same build:

    GET /slots -> [{"id": 0, "n_ctx": 262144, "is_processing": false, "n_prompt_tokens": 120711}]

No site

Links install, modelos, releases.