Pull requests / #332

#332 serve: add a standalone API request monitor

closed · @KadoBOT · 0 コメント · GitHub で見る

BenchmarksServer & APISecurityWindows

本文

The existing Monitor tab shows engine and hardware statistics, but API clients cannot inspect a particular request's input, answer, reasoning, or end-to-end latency. This adds a standalone `/api-monitor` page and request-inspection endpoints, without requiring a chat session.

### Behavior

- Capture parsed OpenAI and Anthropic generation requests, including streamed answers and failures. Show input, output, separate reasoning, non-stream response bodies, status, token usage, engine timings, queue wait, model loading time, first-token latency, and total wall-clock time.
- Retain the newest 100 requests in memory until restart, with each captured text field limited to 262,144 characters and visible truncation flags. Capture limits do not shorten API responses. Request headers, including API authentication credentials, are not recorded.
- Protect `GET /api/requests` and `GET /api/requests?id=<id>` with the existing API-key check. History contains sensitive input/output, so the documentation recommends configuring a key before exposing it on a network.
- Reuse the existing `/load` and `/unload` controls. The page uses relative URLs and the existing UI theme tokens; captured text is rendered with `textContent` rather than HTML.

This PR is independent of the lazy-startup and JSON-response-format contributions. It does not change native inference or claim a performance improvement.

### Validation

- Native Windows: `python -m unittest discover -s serve -p "test_*.py" -v` — 86 tests, OK, 3 optional tests skipped.
- `node --check serve/web/monitor.js`, Python compilation, and `git diff --check` passed.
- Served browser verification against this branch with a CPU-only reloadable mock engine: request selection, input/output/reasoning panes, and load/unload state transitions.
- Regression coverage includes both API dialects and stream modes, generation failures, authenticated history access, capture limits/eviction, and concurrent requests waiting for a cold load. No GPU benchmark was run for this PR.

関連リンク

インストール・モデル・リリースへの站内リンク。