Issues / #980

#980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missing

open · @talisp · 1 评论 · 在 GitHub 查看

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindowsLinux

描述

> **Field report:** a local Qwen3.8-Flash-Next served by Strata, used as the "brain" of an agent (Hermes Agent on Windows 11) that builds and checks 3D scenes in **Unreal Engine 5.8** through the official Unreal MCP. Four days of work (2026-10-01 → 10-05): what broke, why, what fixed it, and what is still missing.
>
> Companion to **#971** (hardware, VM/vGPU, CUDA 12 on driver 550, UD-IQ4_XS with images) and **#804 / #970** (tool call written inside `<think>`). Those details are not repeated here.
>
> Numbers come from the Strata logs, `/metrics`, the Hermes session database and the project's own records. Counts are **VERIFIED** unless marked otherwise. Anything that is only an opinion is marked **(impression)**.

## TL;DR

**What works today:**

- The agent works in a real UE 5.8.3 project (an animation project with a throne room). It:
  - places and poses skinned characters (Mixamo animations);
  - fixes lighting and staging;
  - frames and captures the viewport, and judges the capture with the model's own vision;
  - writes its evidence to a structured log.
- It ran a 5-step experiment queue unattended: **198 tool calls in ~90 min, 7 human messages, none of them a repair**.

**What made the difference:**

1. **Official Qwen sampling on the server.** It had been greedy, which caused runaway loops.
2. **Turning vision on.** We ran ~20 h "blind" without noticing.
3. **The `<think>` tool-call fix** (#970).
4. **Project skills written from the agent's own error log.**
5. **Numeric gates** (projection, bounds, traces) **before any visual judgement.**

**Before → after:**

| | Before | After |
|---|---|---|
| Runaway repetition loops | 3 in ~4.5 h | **0 in the last ~19.5 h** |
| Agent stops after announcing an action | 4 times | **0** (5 recovered by the server) |
| Unreal MCP error rate | ~20–25% | **~8%** on tasks the skills cover |
| Tool calls per human message | 11–17 | **~28** (night run) |

**"20× better" is an impression.** The measured proxies are 2–3×, plus whole failure classes removed (loops, stranded calls) and vision added.
**Still missing:**

- Vision is a second opinion, not a judge.
- Modelling furniture from primitives does not work yet (see "A design failure").
- New MCP toolsets still cost 15–19% errors.
- Long unattended runs need harness discipline.
- The current blocker (the Sequencer stops driving the mesh in the editor) is on the engine side.

---

## How it fits together

```
 Windows 11                                              Linux VM + GPU (see #971)
 ┌──────────────────────────────────────┐   OpenAI API   ┌─────────────────────────────┐
 │ Hermes Agent                         │  (ssh tunnel)  │ Strata server               │
 │  main model ─────────────────────────┼───────────────>│  Qwen3.8-Flash-Next         │
 │  vision_analyze (aux task) ──────────┼───────────────>│  (IQ3_S → UD-IQ4_XS)        │
 │  goal judge: cloud model,            │                │  image encoder on the GPU   │
 │    local Qwen as fallback            │                │  /health, /metrics          │
 │                                      │                └─────────────────────────────┘
 │  MCP client ── HTTP ──> Unreal Editor 5.8.3 (official MCP server)
 │  execute_code / terminal / read_file / patch (local Python tools)
 │  project files on a network drive (rclone SFTP) <── same files ──> VM disk (git only here)
 └──────────────────────────────────────┘
```

**How the agent "sees" the viewport.** Images reach the model only through the auxiliary `vision_analyze` task: a separate, short request with the image and a question. The agent's own turns never carry images.

1. **Frame by numbers.** One `execute_tool_script` call:
   - places the editor camera;
   - projects the 8 corners of the target volume (`WorldPosToScreenCoords`);
   - moves until the target fills ≥12% of the frame;
   - checks occlusion with `trace_world`.

   It returns PASS/FAIL without asking the model anything.
2. **Capture inside Unreal.** `CaptureViewport` (pass `captureTransform`), then write the base64 PNG to disk from Unreal. Only the *path* goes back through MCP: a lit capture is 3–6 M characters, and the client cuts tool results at 2 M.
3. **Decode locally** and check the PNG signature.
4. **Ask the model:** a **neutral** question first ("describe only what is clearly visible in the centre"), then specific *yes / no / can't see* questions.

---

## What went wrong, and how it was fixed

| # | Symptom | Root cause | Fix | Status |
|---|---|---|---|---|
| 1 | "Response dominated by repeated text"; one reply reached **102,833 tokens** | No `sampling` in the Strata config, and Hermes sends none, so the engine was **greedy** | Qwen thinking defaults in the config (temp 1.0, top_p 0.95, top_k 20, presence 0) + `reasoning_budget_tokens: 16000` | ✅ No loop in the last ~19.5 h |
| 2 | Agent announces an action, then stops ("reasoning-only clean stop") | Tool call written inside `<think>` with no `</think>` | #970 | ✅ 4 stops before, 0 after; 5 recovered by the server |
| 3 | Visual checks "ambiguous", errors, or skipped | **Post-install check we skipped:** setup asks "Do you want images?" (default **No**, so a `--yes` install has none). `/health` showed `images: false`, the first `vision_analyze` got HTTP 400, and nothing in the agent turned that into a stop | Answer yes / `--vision gpu`, then check `/health` → `"images": true` and send one test image; later UD-IQ4_XS with images (#971) | ✅ ~20 h of blind work before we noticed |
| 4 | Vision answers looping or timing out (~120 s) | Hermes' auxiliary vision task defaults to **temperature 0.1** | `auxiliary.vision.temperature: 1.0` | ✅ loops / ❓ timeouts stopped later, unexplained |
| 5 | Captures described as "wireframe", contact "ambiguous" | Viewport in wireframe/unlit, or too dark | Capture in **Lit**; reject empty or black frames; "ambiguous" = NOT_EVALUATED | 📋 rule (view mode not settable via MCP) |
| 6 | Files with `N\|` prefixes, a 0-byte state file, `stale_write_blocked`, corrupted `.git/index` | The model did not know the real return shapes of the file tools; failed renames on the network drive | Return shapes documented in `AGENTS.md`; git and script edits only on the VM | 📋 contained |
| 7 | Broken PNGs, context pressure (4 compactions in one session) | Huge MCP results (largest 2,000,083 chars) | Captures via disk; schemas dumped to files and searched | ✅ worked around |
| 8 | **71 errors in 276 MCP calls** in the first sessions | MCP tools hidden behind a meta-tool (JSON quoting broke); real UE 5.8 schemas differ from the model's guesses | Workaround if your model breaks JSON through the meta-tool: `tools.tool_search: off`. Plus a project skill with one rule per observed error | ✅ ~8% on known paths, 15–19% on new toolsets |
| 9 | Agent stops after one item or at the turn limit | `max_iterations=60` | `agent.max_turns: 250`, `/goal` with an explicit stop condition, rules on when a turn may end | ✅ night queue ran to the end |
| 10 | Wrong figure judged; details invented ("filling-in") | Several figures in frame; small-model vision | One target per capture, named by position; numeric gates first | 📋 protocol |
| 11 | Two agent sessions drove the same editor for 12 min | An old session with an active `/goal` resumed after a restart | Stop it, `/goal clear`, discard that window's data | 📋 procedure |
| 12 | The Level Sequence stops driving the mesh in the editor | Editor/engine side (open) | Under investigation | ⏳ The agent refused to mark the recipe "verified" (correct) |

**Row 12, the open Sequencer blocker:**

![Night lab E2: character standing after Sit To Stand](https://github.com/user-attachments/assets/89b8e67f-55a0-40b3-a5e0-7fba6924ee07)

*Night lab E2: after Sit To Stand with KeepState, the centre character ends standing in front of its chair (frame 80). The only time the recipe worked.*

![E5: character still seated, Sequencer not driving the mesh](https://github.com/user-attachments/assets/5b4c218f-d67a-47f1-8a14-4f12bf0e5223)

*E5 after an editor restart: a fresh sequence with the same recipe leaves the character seated at frame 80. The Sequencer stopped driving the mesh (engine/editor side, open).*

<details>
<summary><b>Details on the costliest items (1, 3, 8)</b></summary>

**1. Greedy loops.**
- Hermes sent no sampling fields: checked in 7 saved request dumps.
- Strata keeps the engine default (greedy) when the temperature is absent.
- After the fix, the longest reply was ≤27,752 tokens. Part of that drop comes from the 16k reasoning budget.
- One later loop came right after a **503,090-character** tool result, so huge noisy context may also trigger loops **(impression, one case)**.

**3. Blind for ~20 hours.**
- The base model is multimodal, so everyone assumed it could see.
- The agent judged poses by numbers only and wrote "capture not judged by image".
- Once vision was on, a capture exposed legs going straight through a seat (a component-flag bug) that the numbers had missed.
- **Lesson:** after any server (re)install, check `/health` → `"images": true` and send one synthetic-image probe before trusting a visual step.

**8. MCP schema friction.** Breakdown of the 71 errors:
- ~18 in `execute_tool_script`: wrong module, `run()` not returning a dict, code outside a function;
- 7 × `'calls' is not valid JSON`, through the meta-tool;
- 9 guessed property names;
- the editor paused the MCP for ~60 s after 3 consecutive errors, 9 times.

Schema traps we hit:
- `set_properties.values` must be a **JSON string**; with an object it just returns `false`;
- `CaptureViewport` needs `captureTransform`, although the schema says it is optional.
</details>

---

## What works today

**Flows the agent completes on its own:**

1. **Seat a character with a Mixamo animation.**
   - Import the FBX against the existing skeleton.
   - Set `AnimationSingleNode` + `animationData`, then `bUseRefPoseOnInitAnim=false`.
   - Turn the *component* to face the seat.
   - Read back every property.
   - Check with bounds (seated ≈ [pivot−44, pivot+136.5] cm vs. standing [pivot, pivot+180.5]).
   - Capture and ask vision.

   Five characters were done this way.
2. **Stage background characters** against a wall: a trace from the chest expects ~35–40 cm. Light them with fill lights measured as a figure/ambient luminance *ratio* on the same camera, because absolute values drift.
3. **Unattended lab night.**
   - A JSONL queue of 5 experiments with conditions (E3 only if E2 is blocked, and so on).
   - Limits of 3 strategies × 3 attempts and ~60 min per experiment.
   - One JSONL line per attempt, with 4 gates (structural, geometric, capture, vision), metrics and capture paths.

   Night 1 ran it all in ~90 min and correctly reported a FAIL with evidence.

**Protocols that worked:**

- **Gates before eyes:** structural → geometric → capture → vision. Numbers PASS + image FAIL = FAIL. An empty or timed-out vision answer is NOT_EVALUATED, never PASS.
- **Vision as a second opinion:** a neutral question first, one target per image, named by its position.
- **Operational memory in plain files:**
  - a state file with the next action;
  - a recipe file (21 recipes, each with how it was validated);
  - a failure log (19 entries).

  The agent reads the **state file always**, opens the **recipe for the task at hand**, and searches the **failure log by symptom**. Dumping the whole history every session costs context (one session went through 4 compactions). **(impression)** This is what makes a small local model look much smarter than it is.
- **Safe saves:** only explicit asset paths (`save_assets` with an empty list saves *everything dirty*). Experiments run in a duplicate level.
- **Morning review:** a script packages the night's evidence (summary + thumbnails) for a stronger cloud model to judge. If the cloud judge is unavailable, the night must not stop.

**Where the Unreal skill came from.** No public Unreal skill existed (the hubs have Blender ones).

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。