Pull requests / #837
#837 server: "return_progress": true puts the prompt's PP line on the stream as llama.cpp's prompt_progress
closed · @eliaskg · 0 comments · View on GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quantsWindows
Description

## What
`"return_progress": true` on a streaming `/v1/chat/completions` request adds llama.cpp's `prompt_progress` to the stream: `total`, `cache`, `processed`, `time_ms`, same names and same meaning. A client that already draws a prefill bar for llama.cpp draws one here with no client-side change. Off unless the request asks, as it is in llama.cpp.
The image is [pi-prefill](https://github.com/eliaskg/pi-prefill) running against Strata on this host, needing nothing but its provider id added to a list.
## The numbers were already in the server
The engine writes `PP <position> <prompt_tokens> <ms> <tok/s>` after every prompt chunk (`src/program/generate.cpp`, the protocol note). `serve/server.py` already parses three of those four fields into `engine.progress` and `prefill_tok_s_mean`, the web app draws its prefill bar from them (`serve/web/app.js`, `live.prompt_read` / `live.prompt_total`), and the server window prints them every second while a prompt is read.
The one missing hop is the last. `Service.run` turns a PP line into `("ping", None)`, `openai_chunks` turns that into `None`, and the writer sends it as the SSE comment `: keep-alive`. This fills that comment in. `PP`'s `ms` field was on the wire and discarded, and `RESUME` was parsed into a local and not kept, so those two are now kept as `progress_ms` and `reused`.
## One constant, and why
The batched path never reports the whole prompt: up to `--short-read` tokens (64 by default) are read through the verify windows instead, so its final PP stops short of `pp_total`. Over 14 requests on this host every last PP was **1 or 5 tokens short, never 0**, so `processed >= total` is not the end of the read and a client would watch a bar stop just short of full.
A prompt under one `--prefill` chunk sends exactly one PP, and it is that line. Without the rule such a prompt shows a bar that appears already full. `PP_DONE_TAIL = 64` is `short_read`'s default; a non-default `--short-read` is not read back.
## Measured, two RTX 3090s + 91 GiB, IQ3_S, `--prefill auto` (8192)
A 22,056 token prompt, `return_progress: true`:
```
total=22056 cache=0 processed=8192 time_ms=3728 37.1%
total=22056 cache=0 processed=16384 time_ms=5731 74.3%
```
The step is the chunk, so the bar starts at 37% rather than at 0%. A 42,131 token prompt read in 15 s sends six lines, and a 178,084 token prompt sends about 21. The same 22,056 token prompt sent again reuses the whole prompt from its conversation checkpoint, reads in 0.17 s, and sends **nothing**: the tail rule, working.
## Tests
Four cases in `serve/test_server.py` on the mock engine: one progress chunk per chunk with the exact four fields; nothing at all when the client does not ask, and the keep-alive still there; the tail line dropped at 64, 63 and 1 token short of the end; and `cache` 0 when no `RESUME` came.
`serve.test_server` runs 143, against 139 before. `serve/test_responses.py`'s `test_json_schema_text_format`, and `tools/test_setup_amd / _choices / _golden`, fail the same way on an untouched v0.1.39 checkout on macOS, so they are left as they are.
## Not covered
Chat Completions only. The Responses API and `/v1/messages` have their own ping branches and are untouched. `structured_chunks` passes a progress chunk through instead of buffering it, because it carries no content to validate.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.