Pull requests / #435

#435 setup: resuming a download asks only for the missing space, and a finished .part is not re-requested (#425)

closed · @homeofe · 0 评论 · 在 GitHub 查看

Setup & installAMD / HIPNVIDIA / CUDAModels & quantsLinux

描述

Fixes #425. These are two fixes to resuming an interrupted model download, in two commits.

## 1. The free-space check asks only for what is still missing

**Problem.** When any shard lacked its `.done` mark, the check in `main()` asked for the model's whole `download_gb` again. It ignored finished shards and the `<shard>.part` files that `download()` continues with a `Range` request. The reporter's "~112 GB" for IQ2_XS on a 32 GB PC is exactly 68 (download) + 8 (margin) + 35.5 + 1 (the low-RAM `experts.bin`). The unmodified code prints that same `need ~112 GB` in the new test.

**Change** (`setup.py`):
- New `downloaded_bytes(shards)` (next to `download()`). It adds up the size of every shard that has its `.done` mark, and of the `.part` file for every shard that doesn't. A shard file without a mark isn't counted: a whole one has just been marked by the #173 check above, and a short one is downloaded again into `.part` beside it, so it frees nothing until the end.
- The download term becomes `max(0, download_gb - downloaded / 1e9)`. The 8 GB margin, the 40 GB AVX-512 Q2_0 pack, the image encoder's +1 and the low-RAM arena +1 are unchanged.
- When something is already there, the message says so: `need ~81 GB (not counting the 31.9 GB of IQ2_XS already downloaded)` (the reporter's case, 32 of 68 GB there: 68 - 31.9 + 8 + 36.5). A fresh download prints exactly what it did before.

**Choices**:
- **Shard 2 hard-linked from another size:** once it's linked into this model's folder it has its own mark and takes no new space, so it's subtracted once. One that will only be linked in step 5 is still counted in full, because `os.link` can fail (exFAT, another drive) and setup then downloads it. That case over-asks by one shard, as before.
- **mmproj and `experts.bin`:** still counted as before. Their paths are only resolved later (`find_in(roots)`), the pack can be on another drive than `models_dir`, and a half-written `experts.bin` can't be told apart from a whole one at this point.
- **A server that ignores `Range`:** `download()` restarts that shard, so up to the `.part` size more is needed. The 8 GB margin covers it, as it already covered other estimates.

## 2. A download stopped before its rename finishes without a request

Found while testing the above. If setup stopped after the last byte was written but before `part.replace(dst)`, the next run resumed with `Range: bytes=<size>-`, a range past the end. The server answers **416**, which is an `HTTPError` (an `OSError`), so `download()` printed "download interrupted … retrying in 10 s" **30 times, about 5 minutes**, and only then renamed the complete file. Now the loop stops before asking when `.part` already holds `Content-Length` bytes.

## Tests

The new `tools/test_setup_resume.py` has 11 tests. It runs `setup.main()` through `test_setup_golden.install()`, as `test_setup_risk` does, with `--models-dir` set to a temp folder of small stand-in files. A `float` subclass records the exact `need` that `free_gb()` is compared with.

| test | before | after |
|---|---|---|
| fresh download, 32 GB PC low-RAM / 96 GB PC: `need` unchanged (112.5 / 76), message unchanged | pass | pass |
| shard 1 finished + a `.part` of shard 2: `need` drops by exactly those bytes | **fail** (112.5 ≠ 112.49575) | pass |
| `.part` files that are complete: download term 0, never negative | **fail** (44.5015 ≠ 44.5) | pass |
| whole but unmarked shards are marked first: no download term | pass | pass |
| a short unmarked shard isn't counted | pass | pass |
| the too-little-space message names what is already there | **fail** | pass |
| `downloaded_bytes` unit tests, incl. a hard-linked shard counted once | **error** (no function) | pass |
| a complete `.part`: `download()` makes no GET (HEAD only), renames and marks it | **fail**: 30 × `GET bytes=11-` | pass |

```
python3 -m unittest tools.test_setup_amd tools.test_setup_choices tools.test_setup_draft_vocab \
  tools.test_setup_lowram tools.test_setup_pins tools.test_setup_prompts tools.test_setup_risk \
  tools.test_setup_rope tools.test_setup_unsloth tools.test_setup_resume      # all OK
```

`tools.test_setup_golden` fails on Linux on `main` with or without this change (46 identical failures, the `<EXE>` normalizer). That's fixed separately in #428.

### Test machine

Ubuntu 24.04.4 (kernel 6.8), Python 3.12.3, AMD Ryzen Threadripper 3960X, 128 GB DDR4, RTX 2080 Ti 11 GB. The tests are fully mocked: no network, no downloads, no GPU. They ran beside a live Strata model server on the same PC.

**Not run:** a real interrupted Hugging Face download. A 68 GB model doesn't fit a test, so the 416 behaviour is reproduced with a mocked server.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。