Pull requests / #324

#324 fix(setup): install the engine, model files and packages the checkout was tested with

closed · @alphastorm · 0 コメント · GitHub で見る

Setup & installNVIDIA / CUDAModels & quantsWindowsLinux

本文

Fixes #214.

Setup installed whatever was newest when it ran: the engine from `releases/latest`, the model files, vision encoder and MTP tensors from each Hugging Face repo's `main`, and the Python packages without versions. With this change a checkout installs the same components on every run.

**Change**
- **Engine** (`get_prebuilt`): by default the ready-made engine comes from the release matching the checkout, `releases/download/v<CMakeLists.txt version>/`. Setup falls back to `releases/latest/download/` only on a 404 there, meaning there is no such release or no asset for this system. Any other error (no internet, a rate limit, a 5xx) is handled as before: compile, or keep the installed engine when updating. `--prebuilt` / `STRATA_PREBUILT_URL` is used as given, without a fallback, and `--prebuilt ""` still means no ready-made engine. `PREBUILT_URL` becomes `RELEASES`.
- **The downloaded archive is not kept.** `engine/<asset>` is deleted as soon as it is unpacked, including when setup refuses it. Before, a refused archive (older than `MIN_ENGINE`, or without code for the GPU) stayed behind with its `.done` mark, and every later run reused it instead of downloading again. So the update path's advice to "run this again in a few minutes" (#58) could not work until the zip was deleted by hand, and with two candidate URLs one release's archive would be reused for the other.
- **Hugging Face:** every `hf` / `mmproj_hf` URL in `FAMILIES`, and `REPO` in `tools/mtp_fetch.py`, names a commit instead of `main`. The commits are the four repos' current heads, unchanged since 2026-09-29 or earlier: Qwen GGUF `ed59f92`, Swift `b22d729`, Coder `5348543` and the BF16 checkpoint `de4b8e4`. The files are therefore the ones setup downloads today. The unused `HF` constant is removed.
- **MTP inventory:** `mtp_fetch.fetch` reuses `mtp-inventory.json` only if it was read from the pinned `REPO`. Otherwise it reads the inventory again and drops `tensors/`, because resuming a tensor by its size could join bytes from two commits. Setup runs it only until `mtp/rt/experts.bin` exists, so finished installs are not affected.
- **Python packages:** `requirements.txt` pins the ten packages and their dependencies (16 in all) to exact versions. `setup.py` installs it with `pip install -r`, and the Dockerfile installs the same file.
  - The versions are what an unpinned install picks today: the newest release of each when 0.1.30 came out. numpy is 2.2.6 on Python 3.10, 2.4.6 on 3.11 and 2.5.3 on 3.12 and newer, because numpy 2.5 needs 3.12 and 2.4 needs 3.11.
  - `pip_install` compares the file's lines with `.strata-pip.json`, so a changed pin installs again. An existing venv, whose stamp lists the bare names, installs the pins once on its next setup run.
  - The file has no hashes: `--require-hashes` would need every platform's wheel hashes for each package.

**Behavior changes**
- On Linux, where the releases have no ready-made engine, setup makes one more HEAD request. Before the existing "no ready-made engine" warning it prints "The v0.1.30 release has no ready-made engine for this system: trying the newest release".
- An MTP fetch that was interrupted before this change starts over once (up to ~5 GB).

**Tests** (`python -m unittest tools.test_setup_sources`, no network)
- `Engine`: the checkout's release; the newest release after a 404; a 503 compiles instead of taking another release; `--prebuilt` is taken as given, as is `--prebuilt ""`; a refused archive is not reused.
- `Requirements`: every line is an exact version; a changed pin installs again.
- `HuggingFace`: every URL names a 40-hex commit. An MTP inventory from the pinned commit is reused; one from another commit is read again, and its tensors are dropped.

On `30ec18e` (0.1.30) the new module fails to import (`setup.RELEASES`). With the change it passes (10 tests). Every existing Python test module gives the same result as on `30ec18e`: 158 tests, 5 skipped, with `STRATA_GGUF_PY` set for `tools.test_iq_pack`. This was on Python 3.13.15, macOS arm64. ruff and pyright report nothing new in the changed files.

**Also checked**
- **Stale archive:** the real `get_prebuilt` and `download` ran over a local `--prebuilt` folder that first held a 0.1.0 engine and then a 0.1.30 one. On 0.1.30 the second run prints "Strata engine already downloaded" and refuses the 0.1.0 archive again. With the change it installs 0.1.30.
- **Live releases** (real HEAD requests, download stubbed):
  - a 0.1.30 checkout takes `download/v0.1.30/strata-windows-x64.zip`;
  - a 0.1.27 checkout takes `download/v0.1.27/…`;
  - a 0.1.99 checkout takes `latest/download/…`;
  - the Linux asset falls back and then compiles.
- **Hugging Face:** all 17 files setup can download resolve at the pinned commits (HEAD 200, with `Content-Length` equal to the size in the repo's file tree). So do the BF16 checkpoint's index and an MTP shard. At these commits the Qwen and Coder shard 2 files are identical (same SHA-256), and so are their vision encoders. The shard-2 hardlink and the shared encoder path still hold.
- **requirements.txt:**
  - `pip install --dry-run --only-binary=:all:` installs exactly the pinned set. It ran with each Python's own pip in a fresh venv (3.10 to 3.14, pip 23.0.1 to 26.2.1), for `win_amd64` and for manylinux up to glibc 2.28.
  - `uv pip compile` with each target's markers gives the same set, so no dependency is left unpinned.
  - In fresh venvs on 3.10 and 3.13, `setup.pip_install(REQUIREMENTS)` installs exactly the pins and skips the second call.
  - The Dockerfile's pip step on `ubuntu:24.04` amd64 (the CUDA image's Ubuntu and Python 3.12) installs the file, and the imports work.

Not run: a GPU, a real engine or model download, `START-HERE.bat` on Windows, or a full `docker build`.

関連リンク

インストール・モデル・リリースへの站内リンク。