Pull requests / #122

#122 serve: add llama.cpp-compatible discovery endpoints

closed · merged 2026-09-29 · @hendrikp · 0 comentarios · En GitHub

Server & APIModels & quantsSecurity

Descripción

Fixes #117.

Add read-only metadata endpoints so clients such as Pi can discover the loaded model, context capacity, and thinking/image capabilities without receiving a 404 from `/props`.

- `/models` and `/v1/models` list only the loaded model, including its status, context limit, and input/output modalities.
- `/props` exposes the original chat template, configured generation defaults, context limit, model path, and engine version.
- `/slots` reports the single processing slot and its busy/idle state.
- Metadata endpoints respect the existing API-key setting.
- Discovery does not load, unload, or restart models.

Context capacity reflects the full configured context, not the smaller resident KV window. Image support is reported only when the vision component is loaded. Unconfigured sampling fields are omitted; shared settings take precedence over configuration defaults.

`/props?model=<loaded-model-id>&autoload=false` is supported. Unknown models return 404. When the engine has stopped, model/slot lists are empty and `/props` returns 503.

## Example responses

Examples from a loaded, idle Q2 model with a 262,144-token context and vision disabled. The chat template is shortened for readability and the model path is replaced with a generic example.

### `GET /models` (also `GET /v1/models`)

```json
{
  "object": "list",
  "data": [
    {
      "id": "qwen3.8-flash-next-q2_0",
      "object": "model",
      "status": {"value": "loaded"},
      "meta": {"n_ctx": 262144},
      "architecture": {
        "input_modalities": ["text"],
        "output_modalities": ["text"]
      }
    }
  ]
}
```

### `GET /props`

```json
{
  "default_generation_settings": {
    "n_ctx": 262144,
    "params": {
      "temperature": 1.0,
      "top_p": 0.95,
      "top_k": 20,
      "min_p": 0.0,
      "presence_penalty": 0.0,
      "frequency_penalty": 0.0,
      "repeat_penalty": 1.0,
      "n_predict": -1
    }
  },
  "total_slots": 1,
  "model_alias": "qwen3.8-flash-next-q2_0",
  "chat_template": "[full original Jinja template omitted here]",
  "modalities": {"vision": false},
  "models_autoload": false,
  "is_sleeping": false,
  "model_path": "models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf",
  "build_info": "Strata 0.1.21"
}
```

`n_predict: -1` means there is no fixed output cap beyond the remaining context. With a loaded vision component, `/models` reports `input_modalities: ["text", "image"]` and `/props` reports `modalities.vision: true`.

### `GET /slots`

```json
[
  {
    "id": 0,
    "n_ctx": 262144,
    "is_processing": false
  }
]
```

`is_processing` becomes `true` while the slot is processing a request.

## Validation

- All 30 server tests pass, including the added regression tests.
- Added coverage for model discovery, properties, authentication, vision metadata, generation defaults, build/model information, slot state, and stopped-engine behavior.
- Pi 0.87.1 discovered the loaded model with 262,144-token context and thinking support, then completed a chat request.
- Manual test with pi /models /thinking on/off, behaved correctly - total ctx correctly discovered

En el sitio

Enlaces a install, modelos, releases.