Pull requests / #645

#645 Update the engine from the web app, cautiously

closed · @demetree · 0 comentarios · En GitHub

Setup & installServer & APINVIDIA / CUDASecurityWindows

Descripción

## Please read #670 first

I did not know, when I started this, that `setup.py` already replaces an outdated engine:
`update_installed_engine()` runs on every `START-HERE.bat` (setup.py:3042, 3058), and `UPDATE.bat`
(#475) does it deliberately without starting the model. `get_prebuilt()` even checks the release's own
`archs` against your GPU after unpacking.

So this is a **second way to reach one outcome** — replacing the engine binary — not a missing
capability. #670 asks where you would rather it live, with three options, and says plainly that if the
answer is "keep it to showing the version" then most of this should be closed. I would rather ask than
presume. The engine half of updating already exists; the progress and the backup do not exist anywhere.

## What it actually is

An **Update the engine** card in the About tab. Not a whole-install update: that is `UPDATE.bat`, which
also does `git pull`, the Python packages, and each model's settings and draft subset. This touches the
engine files only, and says so on the card.

Nine steps, and the order is the safety property:

1. **Check the release** — compare `BUILD.json` with `releases/latest`
2. **Check the free disk space**
3. **Download the engine** — to a temp file, against the size the API reported
4. **Inspect the archive** — every member's CRC and path, the version inside, **and whether this card can run it**
5. **Unpack to a staging area** — not over the installed files
6. **Run the staged engine** — `--help`, before the installed one is touched
7. **Back up the installed engine** — `.strata-update/backup-<timestamp>`, three kept
8. **Install the new engine**
9. **Start it and check the version** — read the new `BUILD.json` back

Nothing on disk changes until step 7. From step 8 on, any failure restores the backup and the card says
so, naming the folder.

## One correction to my earlier description of this

I previously claimed the card check ran **before** the download and refused in 0.3–0.5 s instead of
fetching 124 MB. That was wrong, and wrong in an interesting way: it compared the card against the
*installed* engine's `archs`, because the release's `BUILD.json` is inside the zip. In the only
reachable real case — a card below 7.5 running `STRATA_EXPERIMENTAL_SM60=1`, whose engine lists
`sm_61` — it **passed**, and the 124 MB download went ahead anyway. The 0.3–0.5 s figure only existed
because I had hand-edited `BUILD.json` to make it fire.

It now runs at step 4, where the release's own architectures are known, which is where `get_prebuilt()`
does the same check. Still four steps before anything changes, and the message is specific:

> v0.1.38 has no code for this graphics card (compute capability 6.1; it was built for 7.5, 8.6, 8.9, 12)

The cost is the download. The benefit is that it is no longer a guess — and it can now catch the case
that actually bites, a release that *drops* an architecture the installed one had. There is a test
pinning exactly that: a 6.1 card must be **allowed** when the release carries `archs: [61]` and the
installed engine claimed `[75, 86, 89, 120]`.

## Refusals, each measured rather than assumed

| Refused | How |
|---|---|
| A download that stops early | Compared against the size the API reported |
| A corrupt member | `ZipFile.testzip()` before extracting anything |
| A member path outside the destination | Checked before extraction; nothing is written |
| An archive whose `BUILD.json` disagrees with the tag | Compared |
| A release with no code for this GPU | The release's `archs`/`ptx` vs the card, at step 4 |
| A staged engine that will not start | Run before the installed one is touched |
| A downgrade | An older tag is not "nothing to do"; it is refused |
| No build for this platform/backend | Checked in `check()` |
| An update while a request is running or queued | The project's own `unload()` refuses with `busy` |

An unknown card is **not** refused: a wrong refusal costs the user their update, a wrong install costs
them an engine that will not start, and they can always update by hand.

## Progress, not a spinner

Replacing a 113 MB binary and stopping the model is not a "Working..." moment. Every step is listed with
its own state, the download reports its own byte count, and the failure line says **what was restored**
and names the folder. The bar shows the download's byte percent while downloading and the step percent
after it — the first version flashed 100% during the check phase, because that phase has one step and
the run has nine.

State is polled, matching the rest of the panel. An update left running is picked up again when the page
reopens, so the card shows how it ended.

## Guards

Both POST routes carry the same `_own_page` check as `/settings` and `/unload`: JSON only, no foreign
`Origin`, so a cross-site form post cannot trigger an update. `GET /api/update/state` only reads. One
update at a time. And it stops the model first, through the same `unload()` the UI already uses.

## Integrity: a real limit, not an oversight

The release publishes no checksum and `BUILD.json` carries none. What is verified: the download size
against the API, every member's CRC, every member's path, the version inside, the release's own
architectures, and the staged binary running. That catches a truncated download, a corrupt archive, a
wrong artifact and an incompatible one — and it is **not** proof the bytes came from your release.
Closing that needs a checksum in the release process; #670 asks whether you want one.

## Tests

```
python serve/test_update.py          80 checks
python serve/test_update_http.py     45 checks, real HTTP, real handlers
python -m pytest serve              190 passed, 7 skipped
```

No network, no GPU, no 124 MB download — the release JSON, the download and the archive are built
locally. The HTTP test builds a real `Service(MockEngine(...), ByteTokenizer(), ChatTemplate(...))` the
way `serve/test_server.py` does, so a missing attribute is a real bug rather than a gap in the test, and
it no longer depends on the GPU of whichever machine runs it. Cases that must **not** update are asserted
as firmly as the ones that must: truncated download, corrupt archive, unsafe path, version mismatch, an
engine that will not run, an incompatible card, a downgrade, a missing asset, an update while busy.

Rollback was verified over real files by forcing a failure after the apply step: every installed file
came back byte-for-byte (SHA-256 before and after), with the backup left on disk.

**Note on `serve/test_security.py`:** two or three tests there flake on this machine with
`ConnectionAbortedError`, on a pristine checkout with these changes stashed (3/5 runs). A Windows
loopback socket flake, unrelated to this PR.

## Measured on

Windows 11, GTX 1070 (compute capability 6.1), Xeon E5-2680, driver 582.66, CUDA 12.6. Refusal timings
and the download timing are from this machine. The card refusal is specific to a GPU below 7.5, so a
supported card takes the normal path — and on this machine the full success path was exercised against a
local archive holding the engine that does run here, not against your release.

En el sitio

Enlaces a install, modelos, releases.