Pull requests / #350
#350 Chat: a context gauge, per-request timings, and conversation compaction
closed · @orangeswim · 0 commentaires · Sur GitHub
BenchmarksServer & APIModels & quants
Description
# Chat: a context gauge, per-request timings, and conversation compaction Follow-up to the prefill-speed PR (#163 took the Monitor half); this is everything else from that review thread. No server change — `server.py` and `telemetry.py` are untouched. **Context gauge.** `2.4k/200K` with a fill bar next to the sampling button. Amber at 70% of the window, red at 90%. Updates from `usage` after each reply. Passive - it never triggers anything. **Request timings in chat.** A chevron on each reply opens one muted line: `read 25,838 tokens @ 2,165 tok/s · 1,850 reused · TTFT 12.4 s · drafts 214 of 240 accepted`. It is the response's existing `timings` block (llama.cpp's names) - kept in localStorage with the message. The requests table gains a PP t/s column beside decode, computed in the UI from the history fields the server already returns (net of the reused cache share). **Compaction.** The compact button summarizes the older turns: - The newest turns stay verbatim up to a ~5k-token budget - a token budget, not a message count, since six large pastes would defeat the fold; a single message over twice the budget goes to the summarized part instead, however recent - The summary is one hidden non-streaming request (thinking on low, temp 0): seven sections (intent; facts and decisions; verbatim code and artifacts; all user messages; errors and fixes; open items; current thread), told to extract rather than describe and that it may be long - capped at 5k tokens, clamped to the window's remaining room (the server 400s an over-window request), with a warning when the cap truncates it. The prompt lives in its own file (`web/compact-prompt.js`), editable without touching app.js - The kept turns are appended to the summary as an exact verbatim transcript (client-built, never paraphrased), so the payload after any compaction is exactly `[summary] + newer messages` - The summary marker lands at the end of the chat with its stats; only the newest summary can be put back, and undo peels one layer at a time. The page keeps every message, muted above the cut - Long user messages (>10 lines / 800 chars) clamp with a Show-all toggle - display only, the API gets the full text The payload's compaction rules live in `web/compaction.js` (pure; loaded by the app and by `web/test_compaction.mjs` - the cut boundary, chained summaries, undo-after-a-merge, the full-coverage marker shape, corrupt-marker handling; `node serve/web/test_compaction.mjs`), run with the suite through `serve/test_compaction.py` (skips loudly where node is absent). All serve tests pass unchanged. ### AI Tooling Claude-code with GLM-5.3, Strata qwen3.8-flash-next-coder-iq1_m <img width="1124" height="878" alt="image" src="https://github.com/user-attachments/assets/cf7c401a-67c0-4851-820f-9c39c59cbae1" /> <img width="950" height="529" alt="image" src="https://github.com/user-attachments/assets/e31f2967-c7a6-4407-8361-9adc661aaefd" /> <img width="877" height="483" alt="image" src="https://github.com/user-attachments/assets/dbb5bd2e-3ba6-4f7f-9c30-10142225e2a0" /> <img width="854" height="414" alt="image" src="https://github.com/user-attachments/assets/676f3574-b575-47ec-b3ea-14a1dd386ecf" />
Sur le site
Liens install, modèles, releases.