Issues / #756

#756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)

closed · @ikura2024 · 2 comments · View on GitHub

Setup & installServer & APINVIDIA / CUDAModels & quantsDocumentationWindows

Description

**Title:** Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)

**Environment**
- Windows 11 26H2 (OS build 26300.9550), RTX 4070 12GB, RAM 64GB
- Installed by cloning the repository and running START-HERE.bat. `git pull` returns "Already up to date"
- Setup output: "Strata is updated (engine 0.1.38)", model qwen3.8-flash-next-iq3_s is up to date
- Settings: KV cache 8-bit (default), experimental speed projection off

**What is happening**
- When launching with START-HERE.bat and opening the browser, the llama.cpp UI (with "llama-ui" in the top left) is displayed instead of the Strata web app (chat with the graph-based Monitor shown in the README).

**Expected behavior**
- The Strata web app (chat + Monitor with graphs, as shown in the README) opens, or I can find where the Monitor is.

**Additional context**
- Opened URL: http://127.0.0.1:8080/ (the URL the startup log tells me to open)
- The startup log does not print an "engine 0.1.38" line; the version appears only in the setup output.

- Question: is the llama.cpp UI the intended page at `/`, and where is the Monitor with graphs?

**Startup log**

```
Strata - Qwen3.8-Flash-Next on a normal PC (a GPU + system RAM + CPU)
  [ok] GPU: GPU 0 (NVIDIA GeForce RTX 4070, 12 GB)

  ----------------------------------------------------------------------------------------------------
  Starting qwen3.8-flash-next-iq3_s: it loads about 55 GB into RAM and locks part of it for the GPU.
  While it does, YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES (longer the first time after a
  restart). That is normal: please wait and don't close this window - the browser opens when it is ready.
  Later, closing this window stops the model.
  ----------------------------------------------------------------------------------------------------
  Settings (strata-iq3_s.json): --expert-cache auto --prefill auto --spec 4 --max-context 131072 --kv
    int8 --vision --vram-reserve-mib 700 --pcie-frac 0.20 --spec-min-p 0.70; server 127.0.0.1:8080, gpu
    0
loading the vision encoder ...
loading the model (the first start takes a minute or two) ...
[strata] starting the engine: reading the model's weights ...
[strata] loading the experts into RAM (about 55 GB) and locking part of them for the GPU.
         YOUR PC CAN BE SLOW OR STOP RESPONDING FOR 1-3 MINUTES NOW - this is normal.
         Please wait and don't close this window; the browser opens when it is ready.
[strata] experts loaded: 46.84 GiB at 6.17 GiB/s (20 s so far)
[strata] filling the GPU's expert cache (1026 experts, 1.98 GiB of VRAM) ...
[strata] almost ready ...
ready: http://127.0.0.1:8080/v1  (OpenAI: /v1/chat/completions, Anthropic: /v1/messages, context 131072 tokens, images on)
       open http://127.0.0.1:8080/ in a browser to chat; close this window to stop the model
```

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.