Pull requests / #908

#908 setup: ModelScope as a download source (--source), used when Hugging Face's files do not answer

closed · @1872183316 · 0 コメント · GitHub で見る

Setup & installNVIDIA / CUDAModels & quants

本文

setup can download every file from [ModelScope](https://www.modelscope.cn): `--source auto|modelscope|huggingface`
(or `STRATA_SOURCE`). `auto` (the default) keeps Hugging Face; it takes ModelScope only when a model file on
huggingface.co does not answer, which is the usual case in mainland China. `HF_ENDPOINT` (#495) still means Hugging
Face (a mirror chosen on purpose).

## What is on ModelScope

Every repository setup downloads from exists there under the same name with the same paths and sizes (checked
2026-10-05 through ModelScope's file API): `ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, `...-Coder-GGUF`,
`ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, `unsloth/Qwen3.8-Flash-Next-GGUF` and `Qwen/Qwen3.8-Flash-Next`
(the MTP tensors).

## How it is checked

ModelScope serves a repository's current files, not a pinned revision. So:

- `download()` maps a Hugging Face file of a known repository (`HF_REVISIONS`) to ModelScope's URL and, when done,
  checks the file against the SHA-256 ModelScope publishes for it (`verify_sha256`, kept in the `.done` mark as for
  the Unsloth files). A wrong file is deleted and setup stops, as before.
- `tools/mtp_fetch.py` with `STRATA_SOURCE=modelscope` reads the checkpoint from ModelScope; `REPO` and `PINNED` are
  then the same URL, so every MTP tensor is still checked against the pinned revision's SHA-256 (#327).
- When ModelScope does not answer for a file, that file comes from `HF_ENDPOINT` / Hugging Face as before.

## Why a file and not the site

From a PC in mainland China on 2026-10-05: `huggingface.co` itself answered, but a file's redirect to its CDN did not
(connection reset), so `auto` asks for a model file (a HEAD request, following the redirect, 10 s). ModelScope's API
answers HEAD with 404, so its side is asked for a file too, the same request `download()` makes.

## Measured on that PC

- `STRATA_SOURCE=auto` picked ModelScope (in under a second) and `download()` fetched and verified a file;
  `STRATA_SOURCE=huggingface` stopped with "cannot reach huggingface.co" as before.
- Earlier the same day, with the same URLs: the Q2_0 shards and IQ3_S's first shard at 11-14 MB/s, all matching
  ModelScope's SHA-256, and Q2_0's shard 1 matching the copy from hf-mirror.com byte for byte; the MTP tensors from
  `Qwen/Qwen3.8-Flash-Next` all matched the pinned hashes (31 tensors, 5.2 GB, ~3.3 MB/s).

## Tests

`tools/test_setup_modelscope.py` (8, no network: a local HTTP server laid out as ModelScope): the source choice
(explicit, `HF_ENDPOINT`, `auto` with either side answering), the URL mapping, a right hash kept, a wrong one deleted,
no published hash, and the fall back to Hugging Face. `tools/test_setup_pins.py`'s gone-revision test now sets
`STRATA_SOURCE=huggingface`, since it counts Hugging Face requests and `auto` probes first. All `tools/test_setup_*.py`
pass.

## Not done

A whole install through setup from ModelScope (the files were fetched by these functions and by hand, not by one setup
run); the ready-made engine and llama.cpp's zip still come from GitHub.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

関連リンク

インストール・モデル・リリースへの站内リンク。