Pull requests / #350

#350 Chat: a context gauge, per-request timings, and conversation compaction

closed · @orangeswim · 0 Kommentare · Auf GitHub

BenchmarksServer & APIModels & quants

Beschreibung

# Chat: a context gauge, per-request timings, and conversation compaction

Follow-up to the prefill-speed PR (#163 took the Monitor half); this is everything
else from that review thread. No server change — `server.py` and `telemetry.py`
are untouched.

**Context gauge.** `2.4k/200K` with a fill bar next to the sampling button. Amber
at 70% of the window, red at 90%. Updates from `usage` after each reply. Passive -
it never triggers anything.

**Request timings in chat.** A chevron on each reply opens one muted line:
`read 25,838 tokens @ 2,165 tok/s · 1,850 reused · TTFT 12.4 s · drafts 214 of 240
accepted`. It is the response's existing `timings` block (llama.cpp's names) -
kept in localStorage with the message. The requests table gains a PP t/s column
beside decode, computed in the UI from the history fields the server already
returns (net of the reused cache share).

**Compaction.** The compact button summarizes the older turns:

- The newest turns stay verbatim up to a ~5k-token budget - a token budget, not a
  message count, since six large pastes would defeat the fold; a single message
  over twice the budget goes to the summarized part instead, however recent
- The summary is one hidden non-streaming request (thinking on low, temp 0):
  seven sections (intent; facts and decisions; verbatim code and artifacts; all
  user messages; errors and fixes; open items; current thread), told to extract
  rather than describe and that it may be long - capped at 5k tokens, clamped to
  the window's remaining room (the server 400s an over-window request), with a
  warning when the cap truncates it. The prompt lives in its own file
  (`web/compact-prompt.js`), editable without touching app.js
- The kept turns are appended to the summary as an exact verbatim transcript
  (client-built, never paraphrased), so the payload after any compaction is
  exactly `[summary] + newer messages`
- The summary marker lands at the end of the chat with its stats; only the
  newest summary can be put back, and undo peels one layer at a time. The page
  keeps every message, muted above the cut
- Long user messages (>10 lines / 800 chars) clamp with a Show-all toggle -
  display only, the API gets the full text

The payload's compaction rules live in `web/compaction.js` (pure; loaded by the
app and by `web/test_compaction.mjs` - the cut boundary, chained summaries,
undo-after-a-merge, the full-coverage marker shape, corrupt-marker handling;
`node serve/web/test_compaction.mjs`), run with the suite through
`serve/test_compaction.py` (skips loudly where node is absent). All serve tests
pass unchanged.


### AI Tooling
Claude-code with GLM-5.3, Strata qwen3.8-flash-next-coder-iq1_m

<img width="1124" height="878" alt="image" src="https://github.com/user-attachments/assets/cf7c401a-67c0-4851-820f-9c39c59cbae1" />
<img width="950" height="529" alt="image" src="https://github.com/user-attachments/assets/e31f2967-c7a6-4407-8361-9adc661aaefd" />
<img width="877" height="483" alt="image" src="https://github.com/user-attachments/assets/dbb5bd2e-3ba6-4f7f-9c30-10142225e2a0" />
<img width="854" height="414" alt="image" src="https://github.com/user-attachments/assets/676f3574-b575-47ec-b3ea-14a1dd386ecf" />

Mehr auf der Site

Links zu Install, Modellen, Releases.