Issues / #756
#756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)
closed · @ikura2024 · 2 comments · View on GitHub
Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows
Description
**Title:** Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)
**Environment**
- Windows 11 26H2 (OS build 26300.9550), RTX 4070 12GB, RAM 64GB
- Installed by cloning the repository and running START-HERE.bat. `git pull` returns "Already up to date"
- Setup output: "Strata is updated (engine 0.1.38)", model qwen3.8-flash-next-iq3_s is up to date
- Settings: KV cache 8-bit (default), experimental speed projection off
**What is happening**
- When launching with START-HERE.bat and opening the browser, the llama.cpp UI (with "llama-ui" in the top left) is displayed instead of the Strata web app (chat with the graph-based Monitor shown in the README).
**Expected behavior**
- The Strata web app (chat + Monitor with graphs, as shown in the README) opens, or I can find where the Monitor is.
**Additional context**
- Opened URL: http://127.0.0.1:8080/ (the URL the startup log tells me to open)
- The startup log does not print an "engine 0.1.38" line; the version appears only in the setup output.
- Question: is the llama.cpp UI the intended page at `/`, and where is the Monitor with graphs?
**Startup log**
```
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
[ok] GPU: GPU 0 (NVIDIA GeForce RTX 4070, 12 GB)
----------------------------------------------------------------------------------------------------
Starting qwen3.8-flash-next-iq3_s: it loads about 55 GB into RAM and locks part of it for the GPU.
While it does, YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES (longer the first time after a
restart). That is normal: please wait and don't close this window - the browser opens when it is ready.
Later, closing this window stops the model.
----------------------------------------------------------------------------------------------------
Settings (strata-iq3_s.json): --expert-cache auto --prefill auto --spec 4 --max-context 131072 --kv
int8 --vision --vram-reserve-mib 700 --pcie-frac 0.20 --spec-min-p 0.70; server 127.0.0.1:8080, gpu
0
loading the vision encoder ...
loading the model (the first start takes a minute or two) ...
[strata] starting the engine: reading the model's weights ...
[strata] loading the experts into RAM (about 55 GB) and locking part of them for the GPU.
YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES NOW - this is normal.
Please wait and don't close this window; the browser opens when it is ready.
[strata] experts loaded: 46.84 GiB at 6.17 GiB/s (20 s so far)
[strata] filling the GPU's expert cache (1026 experts, 1.98 GiB of VRAM) ...
[strata] almost ready ...
ready: http://127.0.0.1:8080/v1 (OpenAI: /v1/chat/completions, Anthropic: /v1/messages, context 131072 tokens, images on)
open http://127.0.0.1:8080/ in a browser to chat; close this window to stop the model
```Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.