Pull requests / #12

#12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…

closed · @chynggi · 0 comments · View on GitHub

BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsSecurityDocumentation

Description

… one-click setup

- setup.py: three uncensored families with their own sizes/shard counts; OrcaRouter is gated - the setup says so and takes a Hugging Face token (HF_TOKEN, hf auth login, or pasted), checked before the download; extra disk for preparing them
- iq_pack.py: layers split across shards -> experts.bin; quantized tensors the engine cannot serve natively (hc_*, ssm_alpha/beta, ple_value...) -> BF16 in dense.bin; native_experts.txt written last so a stopped run leaves no pack
- ple_key_bf16.py: rewrites the shard holding blk.1.ple_key once with the key as BF16 when it is quantized to anything but Q2_0 (still a valid, contiguous GGUF)
- README: the uncensored section

Verified: OrcaRouter IQ2_M downloaded, prepared and served on an RTX 3060 12 GB + i5-10400 + 64 GB RAM, ~21 tok/s; the original Qwen/Swift files are unaffected.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.