Pull requests / #177
#177 Web UI: prefill speeds, request timings, a context gauge, chat compaction
closed · draft · @orangeswim · 0 commentaires · Sur GitHub
BenchmarksServer & APIModels & quants
Description
# Web UI: prefill speeds, request timings, a context gauge, chat compaction No engine changes. This UI change exposes some metrics already tracked. **Prefill speeds.** `_prompt_tok_s()` in server.py (the cumulative mean, not a window - the engine reports prompt progress per 8k chunk, so a window rate swings between 0 and the burst rate), a `prompt_tok_s` telemetry series, and `prompt_tok_s` per request in `/metrics` history. In the Monitor: a Prefill card, the rate on the "Reading prompt" line, `read N @ X tok/s` on the last-request line, a Read t/s column in the requests table. **Request details in chat.** A chevron on each reply opens one line: `read 25,838 @ 2,165 tok/s · 1,850 reused · TTFT 12.4 s · drafts 214 of 240 accepted`. It's the response's `timings` (llama.cpp names), kept in localStorage with the message. **Tokens in/out.** A Monitor card from the server's running totals: everything read vs written since the server started (`1.2M / 340k`, cache reuse excluded from the in side), with the request count and reused tokens as its sub-line. **Context gauge.** `2.4k/200K` with a fill bar next to the sampling button. Amber at 70% of the window, red at 90%. Updates from `usage` after each reply. Passive - it does not trigger anything. **Compaction.** The layers button folds the older turns into a summary; the newest turns stay verbatim up to a ~5k-token budget, the API payload becomes the summary plus everything after the cut, and the page keeps every message (older turns muted above the cut). The summary is one hidden non-streaming request (thinking off, temp 0, 1600 tokens max) with a structured prompt: intent, facts and decisions, verbatim code, all user messages, errors and fixes, open items, current thread. The fold marker lands at the end of the chat with its stats; only the newest marker carries a "put back into the context" undo (an earlier one could not restore anything - a later fold still stands over it), and unfolding peels one layer at a time. Chained folds merge. Long user messages (>10 lines / 800 chars) clamp with a Show-all toggle - display only, the API gets the full text. All 38 serve tests pass unchanged. Tested on a 3090 against 0.1.26 (k8v4 and int8). ### AI Tools Created and reviewed with the help of Claude-Code using GLM-5.3, Strata qwen3.8-flash-next-coder-iq1_m and Codex. <img width="1288" height="848" alt="image" src="https://github.com/user-attachments/assets/16e9a7a2-1781-4973-b2e4-cd80b8d76658" /> <img width="917" height="852" alt="image" src="https://github.com/user-attachments/assets/616a75b4-22d6-450b-afae-cec2055101f4" /> <img width="1010" height="716" alt="image" src="https://github.com/user-attachments/assets/6fced703-f2c8-406f-bd1b-32aa48cf8ad4" />
Sur le site
Liens install, modèles, releases.