Issues / #1626

#1626 [Feature Request]: more of the llama outputs, etc

open · @jeremyoha450 · 0 comments · View on GitHub

Server & API

Description

Posible to add more of the llama outputs, etc? (i've add some myself. but other might like them too)

**Token and template routes**
- `POST /tokenize` is missing upstream. A local patch adds it, plus `/v1/tokenize`, with `with_pieces`.
- `POST /detokenize` is missing upstream. A local patch adds it, plus `/v1/detokenize`.
- `POST /apply-template` is missing upstream. A local patch adds it, plus `/v1/apply-template`, including `response_format` schema text and `input_tokens`.
- There is no llama.cpp-compatible `add_special` / BOS handling on tokenize.
- There is no `/tokenize` of a chat message list without a separate apply-template call.

**Native generation**
- `POST /completion` is missing, including the native stream fields `content`, `stop`, `stop_type`, `tokens_predicted`, `tokens_evaluated`, `tokens_cached`, and `generation_settings`.
- `POST /completions` is missing.
- `POST /infill` is missing.
- `POST /v1/completions` is missing (legacy non-chat completions).

**Embeddings and adapters**
- `POST /embedding` is missing.
- `POST /embeddings` is missing.
- `POST /v1/embeddings` is missing.
- `GET /lora-adapters` is missing.
- `POST /lora-adapters` is missing.

**Slots and metrics**
- `POST /slots/{id}?action=erase` is missing. Save and restore exist.
- Prometheus text on `GET /metrics` is missing. Strata’s `/metrics` is a different JSON document.
- The llama.cpp `timings` object is missing from chat chunks: `prompt_ms`, `predicted_ms`, `prompt_per_token_ms`, `predicted_per_token_ms`, `prompt_per_second`, `predicted_per_second`.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.