Pull requests / #1507
#1507 # Chat: a context gauge, per-request timings, and conversation compaction
open · @orangeswim · 0 comentários · No GitHub
BenchmarksServer & APIModels & quants
Descrição
# Chat: a context gauge, per-request timings, and conversation compaction Re-opening of #350, which GitHub closed when main's history was cleaned up (not a rejection, per the heads-up there). Same change, rebuilt on the new main as three commits, one per feature. No server change. What it adds: - **A context gauge** — a small pill in the composer bar (fill bar + token count, e.g. `2.4k/200K`) showing how much of the context window the conversation uses. Amber at 70% of the window, red at 90%. Updates after each reply; passive. - **Per-request timings, and a PP t/s column** — a chevron on each reply opens one line with that request's numbers: tokens read and their speed, cache reuse, time to first token, drafts accepted (kept with the message). The Monitor's requests table gains a prompt-processing tokens/s column beside decode, computed in the UI from fields the server already returns (the reused cache share excluded; a fully cache-served request shows a dash). Also: the streaming cursor now shows from the moment a message is sent, while the prompt is still being read, not only once tokens arrive. - **Conversation compaction** — a button that asks the model itself to summarize the older turns: the newest turns stay verbatim up to a ~5k-token budget, the summary carries them as an exact client-built transcript, and from then on the API receives the summary plus the newer messages instead of everything. The page keeps every message (the summarized ones are dimmed above the divider, and only the newest summary can be put back — older ones are already contained in the one after them). Long user messages clamp with a Show-all toggle; display only. The payload rules live in `web/compaction.js`, tested by `web/test_compaction.mjs` (run through `serve/test_compaction.py` so `unittest discover` picks them up). Notes for this rebase: - The one conflict (`apiMessages`) was resolved by folding #1392's "the answer being asked for now is not history yet" rule into the new callback form. - The compaction rules were aligned with #1392's error-turn policy: `compaction.js` no longer drops error turns itself — it only decides which messages the summary replaces; the converter decides how a turn is sent (so a turn with no answer text still goes back as an empty assistant turn, as on main). <img width="875" height="819" alt="image" src="https://github.com/user-attachments/assets/93d871e5-4e90-49bc-9d87-3bec3eec91e9" /> <img width="862" height="233" alt="image" src="https://github.com/user-attachments/assets/dacda2a5-b2bd-42a9-a13a-c7f90b979ac7" /> <img width="1109" height="290" alt="image" src="https://github.com/user-attachments/assets/14fc1843-3d6c-4f47-b0dc-c5cb516fdf0c" />
No site
Links install, modelos, releases.