Temas / Setup & install
Setup & install
Auto topic: Setup & install
Issues
- #1495 [Feature]: experimental per-stage KV-grow for a two-GPU layer split at native 262K
- #1488 prefill.cpp: a GCC 10 build fails (std::atomic::wait is a GCC 11 library feature, with no guard and no documented minimum)
- #1487 Windows/CUDA 11.8 Volta integration: focused patches and reproducible handoff
- #1485 [Bug]: update.sh cannot rebuild a locally compiled engine on a PC with a Volta card beside an RTX 40 (nvcc: Unsupported gpu architecture 'compute_70')
- #1482 [Feature]: publish a ready-made Docker image, so users do not compile the engine themselves
- #1475 expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)
- #1474 hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
- #1469 0.1.40.2: decode on a Tesla P40 (sm_61) is about half of 0.1.40's; bisected to 2e4ddf6, and restoring __restrict__ in pdl.hpp (STRATA_PDL_RESTRICT) gives back most of it
- #1463 MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)
- #1453 Live agent workload on RX 7900 XTX (24 GB) + 30 GB RAM — 70–74 tok/s, 91–96% expert cache hits
- #1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combine
- #1445 vision encoder did not start on V100-SXM2-32GB
- #1444 can kvcache be saved and reused on disk?
- #1442 Feature request: support chat title generating from Xcode's agent
- #1440 [Bug]: [SYCL] 0.1.40.2 on 2x Arc Pro B70: --layer-split fails three ways (#1054's fix from #1111 is not in main, plus two new split regressions)
- #1428 --layer-split auto on a P40 + RTX 3070 (3070 first) flips from K=43 to K=47 with --kv int8 or --vram-reserve-mib 600, and prefill drops from 458 to 111-113 tok/s (0.1.40.2)
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying
- #1404 [Feature]: run the image encoder on a second PC
- #1397 SYCL port on Windows/OpenCL: 0.1.40-sycl decodes at 22.3 tok/s with --spec 2 where 0.1.39-sycl does 53.6 (and short runs vary 40%)
- #1389 RX 7900 XTX (gfx1100): 925 -> 2,503 tok/s prompt read (+171%) - the method, and the 26% cliff below full lend coverage
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by default
- #1375 Strata on an RX 6600 XT (gfx1032, 8 GB) and a Xeon E5-2678 v3 with DDR3
- #1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it
- #1353 stager-transient-regression
- #1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stack
- #1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)
- #1344 Server: Multiple API key(llamacpp format) will seen as one long key
- #1337 setup --calibrate loses the measurements it already has when a later engine start fails
- #1322 setup re-run with --vision gpu does not add --vision to an existing config's args
- #1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)
- #1308 Traitement de l'eau et rondelle de roue : les solutions industrielles de Tecnoter Group
- #1285 Dual RTX 5090: Strata 0.1.40 production decode at 243 tok/s and synthetic peer-prefill results
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1276 UPDATE.bat cannot update existing checkouts after the main history rewrite
- #1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speed
- #1261 461
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
- #1244 Docker v0.1.40: persisted /data/config edits ignored on restart; stale /opt/strata config used (worked in v0.1.39)
- #1239 v0.1.40 two-GPU tester results on a P40 + RTX 3070: 150-request soak clean, --pipeline-windows 2 +9% / +15% decode in 5 of 5 pairs
- #1238 0.1.40: --layer-split auto with STRATA_STAGE_TRIM=1 still picks K=2 on a P40 + RTX 3070; 572b607 fixes it but is not in the release
- #1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32
- #1233 ``` Windows + Strix Halo (gfx1151, Radeon 8060S): UD-IQ4_XS runs end-to-end — decode flat 29.5-30.4 tok/s at 1K-32K, but a GPU driver timeout (TDR) during a cold 32K prefill ```
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1212 Review: shared MTP head fix — sibling Q4 state omission and test configuration gaps
- #1211 Bench: --ple-io direct vs. --ple-io ram on system with sufficient headroom of RAM
- #1208 HIP layer split / gfx1030: batched prompt path produces NaN logits (output "!!!!") unless STRATA_HIP_PROMPT_F16=1 and the gfx1030 card is the first stage
- #1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???
- #1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores it
- #1191 Automatic support of other variants and quants
- #1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)
- #1169 k8v4 + --kv-resident: the batched KV append (#783) skips the host mirror for hybrid — parking and KV streaming serve stale bytes (cross-conversation contamination on 0.1.40)
- #1155 [Windows / AMD] --vision cpu is still unreachable: hip_vision()'s WIN gate, find_vcvars()'s VS-18 range, and no encoder in the ready-made engine
- #1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)
- #1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
- #1135 k8v4: deterministic token repetition on sequential generation (RTX 5090, 0.1.40) — byte-identical across two quants
- #1129 New Feature - Add option so pick the default model settings: --thinking --instruct
- #1094 --layer-split auto fails on dual GPUs for 512k context
- #1085 Performance regression after updating to v0.1.40 - RTX 5090
- #1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40
- #1071 vmm.cpp: the CUDA 12 engine does not build with a toolkit older than 12.5 (missing CUDART_VERSION guard)
- #1070 New Feature - add option to change location of model
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
- #1059 serve: a request the engine rejects with ERR waits the full 300 s before failing (STOP + _drain_control("DONE") after an ERR), holding the control lock
- #1053 A reply that ends inside `<think>` (no `</think>`) comes back empty with `finish_reason: "stop"`, and the agent stops — close the thinking and continue once
- #1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)
- #1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
- #1018 [gfx906] gr_up_fast_kernel<GrMulti> aborts the engine with HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION on long-context chats (intermittent; the non-fast path is stable)
- #1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)
- #1005 CRLF line-break tokens (`\r\n`, ids 317/845/8301) get far too little probability; everything else matches llama.cpp
- #990 MCP install planning rejects supported AMD CPU vision (and retains the old Windows HIP restriction)
- #985 Also allow Visual Studio 2026 to build from source
- #984 reasoning_budget_tokens is ignored?
- #980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missing
- #979 HIP: a gfx1201 hipBLASLt table for 1.2.1 (100201, ROCm 7.2.0) makes the prompt 6-17% slower than no table
- #977 --check prints [ok] for below-floor RAM and a misleading 'This PC can run Strata' verdict
- #976 serve: structured-output failure path returns error: null instead of 502 structured_output_failed (test_responses)
- #975 setup.get_prebuilt_hip() zip-unpack branch returns None on Linux (tools/test_setup_amd.py)
- #973 Golden drift: all 46 test_enter_for_every_question subtests fail on main (tools/test_setup_golden.py)
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
- #968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly
- #967 Unsloth ud-q4_k_xl - no vision?
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
- #959 ./setup.sh fails on Gentoo Linux
- #953 Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decode
- #952 BrokePipeError when terminating Strata on Linux
- #942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39
- #938 rx 7600xt 16gb support
- #937 `[gfx1030] HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION in gdn_step_commit_kernel after a few long fresh prefills`
- #932 Dual GPU support
- #929 `--batch` with `--layer-split` exits ("verify batch: layer 34 never rang") with a 3070 first and a P40 second; works with the P40 first (CUDA, v0.1.39)
- #925 Setup: show progress in the browser from the first second (status page with steps, speed and ETA)
- #922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runs
- #918 Windows: Radeon 8060S (gfx1151) builds, selftests and passes the HIP ctest on top of #895
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
- #915 Windows AMD report: RX 6800M (gfx1031, laptop) works with a self-built HIP engine — please add 0x73DF / gfx1031 to the Windows path
- #906 Older PC (Haswell, DDR3, PCIe 3.0 x8): +15% decode on Q2_0 and IQ3_S from #706, #764 and a busier adaptive tier
- #892 Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works
- #890 HIP: "parallel": 2 + --layer-split on 2x R9700 exits (code 1) after "captured the batch window over slots 0,1"
- #884 HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)
- #881 [Windows / AMD] RX 6800M (gfx1031) runs with a self-built HIP engine — plus 2 Windows bugs found (packaging encoding, vision `vcvars`)
- #871 Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)
- #870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
- #841 not stable start with 3+ cards
- #840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaningless
- #836 Milestone: Add Apple Silicon Cross-Platform Support
- #830 Let a one-shot request skip the prompt-cache work (~45 ms per call that is never reused)
- #816 0.1.39: decode on an RX 6800 (gfx1030, Windows) is 12-22% slower than 0.1.38; STRATA_SH_STREAM=0 removes it
- #806 Strange behavior for x2 RTX 5060 Ti (16Gb each) and UD-Q4_K_XL
- #803 Perplexity ~7-10% above llama.cpp on the same GGUF (Qwen3.8-Flash-Next), also at 2K context, every version tested
- #798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`
- #775 Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile
- #767 --vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL grounding
- #756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)
- #754 Anthropic API: tool calls intermittently emitted inside `thinking` and returned as `end_turn`
- #747 Feature request: multi-conversation parking on multi-GPU layer-split
- #729 KV int8 vs fp16 KLD compare at 240K context
- #728 Infinite loop issue
- #725 Issue: API key provided via --api-key flag is not applied in the UI
- #703 [AMD HIP] RX 6750 XT (gfx1031) works in heterogeneous layer split; long multimodal prompts can crash the engine
- #702 [AMD HIP] RX 9060 XT (gfx1200) Repeating output
- #694 fit_max_tokens issue, context problem with agentic coding.
- #692 Feature Request: Intel Arc A770 Support
- #691 Windows: generation slows when the server window is minimized (EcoQoS moves threads to E-cores) — fix included
- #690 Layer split: native prefill kernels fail on the second GPU (no kernel image available), regardless of which card it is
- #670 Keeping the old engine when an update succeeds but turns out bad - and a question about where the web app should call it
- #669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
- #657 Error while starting - PLE table size mismatch
- #654 HIP (gfx1200) RuntimeError Exception Code: 0xC0000005
- #647 Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS works on double Vega 20 GPU (Radeon Pro VII 16gb each).
- #642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)
- #641 AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64): a working port, numbers, and PRs to upstream it
- #629 Re-running setup with a different --context silently discards mcp_servers, mcp and sampling from the run config
- #625 Vision (CPU path): offer a Q8_0 mmproj and a higher image-token cap
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #612 Possible to support gfx1151?
- #609 --no-browser flag?
- #607 Offer to donate AI inference and contribution work to Strata
- #605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)
- #604 Resident expert placement on a layer split, for mixed or older GPUs: where I think it helps, plus some doc notes, and an offer to test
- #602 Measured numbers: Qwen3.8-Flash-Next on 2× RTX 3060 (Q2_0 vs IQ3_S), plus a local reproduction of #75
- #601 Linux build fails on Ubuntu 26.04 (glibc 2.43) with CUDA 12.9 (nvcc host-math conflict)
- #597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090
- #590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 W
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)
- #579 AMD/HIP RX 9070 XT (gfx1201) stalls during batched prompt ingestion on ROCm 10.0
- #566 Proposal: extend existing calibration to Linux HIP
- #564 Gui settings
- #560 AMD/Linux desktop: default VRAM reserve (700 MiB) lets the driver evict ~24 GB to RAM, OOM kills kwin; --vram-reserve-mib 3072 fixes it
- #551 Wrong UI on Windows?
- #549 Update.sh fails to update with a KeyError: 'args' error
- #542 sm_120 + IQ packs: ~4x prefill regression 0.1.34-0.1.36 when libcudart's ABI is mismatched - the #420 fits() gate silently disables MMQ (smpbo reads 1)
- #534 Performance, RTX3060 vs RX7800XT
- #530 high/xhigh reasoning_effort can silently run to max_tokens with empty content when reasoning_budget_tokens isn't set
- #517 new error `raise RuntimeError("the engine exited before it was ready" + (f" (see {log})" if log else "") +`
- #515 Feature request: Intel Arc B390 / Panther Lake iGPU support on Windows (64 GB shared RAM)
- #507 Naive question about context window
- #506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)
Pull requests
- #1499 core: write reduced helper rows directly to mapped output
- #1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegister
- #1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)
- #1490 Add a workflow that builds and pushes the Docker image on a version tag
- #1486 feat: introduce Dockerfile.rocm-stable
- #1476 Update GLM port to Project Maya v1.0.4
- #1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K
- #1458 helper cache: configurable VRAM allowance with allocation-floor checks
- #1449 Add Metal backend support for Strata on macOS
- #1443 Added benchmarks from my setup
- #1438 perf(peer): add opt-in byte-capacity-aware expert placement
- #1437 perf(cuda): add opt-in shape-tuned sm75 interleaved verify
- #1436 fix(serve): isolate CPU cores for independent Linux replicas
- #1434 telemetry: GPU load, VRAM, temperature and power for AMD cards on Windows (#1380)
- #1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)
- #1420 Add optional GLM-5.3-Flash support from Project Maya
- #1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer split
- #1405 tools: add configurable launches and reuse local model downloads
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1401 V100 (sm_70): prompt experts on FP16 tensor cores (+10% prompt) and three existing decode kernels as defaults (-6.8% GPU per window)
- #1398 docs: keeping Codex CLI off the network with Strata
- #1395 hip: a hipBLASLt tuning table for gfx1150 (Radeon 890M), and gfx1151's exact speed switches on gfx1150
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
- #1388 hip: gfx1151 hipBLASLt tuning table for setup's ROCm 7.14.0a20260608 (hipBLASLt 1.4.0)
- #1383 generate: say that CUDA_LAUNCH_BLOCKING=1 hangs the verify windows (#1341 #964)
- #1378 bench: community report, RX 6600 8 GB (gfx1032), EPYC 7232P, IQ3_XXS 128K on Linux
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)
- #1359 setup: link the engine against the toolkit whose nvcc builds it (CUDAToolkit_ROOT)
- #1358 build: vmm.cpp builds with a CUDA toolkit older than 12.5 again (#1071)
- #1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4
- #1332 calibrate: the PCIe share sweep reaches 1.0, where no expert goes to the CPU pool
- #1326 serve: support IPv6 bind addresses
- #1320 test: qualify existing exact HIP paths on gfx906
- #1319 serve: add opt-in bounded Responses history and continuation
- #1318 docs(AMD_HIP): my R9700 numbers for hipBLASLt 1.4.1 and STRATA_HIP_WMMA
- #1313 hipBLASLt table for gfx1201 at 100200 (R9700 calibrated, +15% prefill)
- #1309 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
- #1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #1287 feat(vision): support Windows HIP image encoding
- #1283 prefill: bytes_needed sizes the MoE buffers for the layout init picks (with or without an expert source)
- #1282 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
- #1278 HIP: run the K-quant MMQ test on HIP builds; UD-Q4_K_XL measured on an R9700
- #1274 setup: --thinking / --instruct set the sampling every client gets (#1129)
- #1268 --batch-mtp: the first slot admission ends the engine ("unsupported native MMVQ GGML type")
- #1266 Add the Update the engine card
- #1265 Add the engine updater: verify, stage, all-or-nothing swap, rollback
- #1259 setup: --draft-vocab it - the English/code subset plus the tokens of Italian text
- #1251 --peer-device prompt share without P2P: host route, per-token sums, FP16 transfers (2.1-2.5x prompts on a no-P2P pair)
- #1249 Pipelined batch: two requests run at half speed (pad rows in group windows, unbalanced slot choice)
- #1245 bench: UD-IQ4_XS first NVIDIA measurement (RTX 5090 Laptop, 24 GB)
- #1242 Isolate concurrent image positions and batch-MTP draft paths
- #1241 serve: the Responses API reads Codex's additional_tools input items as tools (#782)
- #1227 mcp: the install plan follows setup's AMD rules - Windows AMD and vision=cpu on Linux are planned (#990)
- #1224 serve: the Monitor's AMD GPU readings on Windows through ADLX (load, VRAM, temperature, power)
- #1219 setup: a Spanish draft subset (--draft-vocab es); Spanish answers dra…
- #1218 Check the ready-made engine against GitHub's SHA-256 before installing it
- #1216 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens (reland)
- #1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoring
- #1205 update.sh: a rewritten history is not the user's fault, and the messa…
- #1202 dashboard: Serve the web page under an optional dashboard path
- #1198 engine: document --spec and --spec-min-p in --help
- #1197 docs: --calibrate loads the model more than once, and starts it when …
- #1192 Community benchmark: 4x Tesla P100 16 GB (Pascal), IQ3_XXS, CUDA 12 engine
- #1186 setup: --draft-vocab es - the English/code subset plus the tokens of Spanish text
- #1185 verify: a window graph that finds no VRAM frees older batch slot layouts' graphs instead of failing (#997)
- #1176 strata_pack.py build --force: unmake the old pack before writing, so a build that stops part-way is an unfinished build, not a pack (#634 follow-up)
- #1173 bench: community report, RX 7900 XTX on Windows 11, IQ3_S, engine 0.1.40
- #1164 conversation cache: a shared prefix is copied out of a parked conversation instead of taking it over, so the parked one keeps its history
- #1162 Verify the downloaded engine against GitHub's SHA-256 (replaces #645)
- #1159 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1156 setup: reach the CPU image encoder on Windows + AMD (#1155, #881)
- #1152 Fix Windows AMD telemetry and add optional Arabic UI
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
- #1132 docs: running Strata behind llama-swap
- #1131 generate: the prompt loan keeps its residency bookkeeping without the token graph (#796 part B)
- #1127 hip: rebased support for Radeon gfx1151 APUs
- #1126 Harden server, installer and engine against high-risk bugs
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)
- #1119 Add local PDF, Word and Excel attachments with screenshot previews
- #1111 Intel arc 0.1.40
- #1110 fix(mtp): --batch-mtp exits on the first admission with a draft-vocab head (slot drafters never copy its type)
- #1107 prefill: gather experts in groups on short prompts too
- #1100 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)
- #1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)
- #1098 prefill/gdn: the Volta prompt recurrence chain-split (+7-15%) and the opt-in chunked path
- #1097 V100: bit-exact GEMV / MoE operator fusions (experimental build only)
- #1096 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.2% e2e)
- #1092 launcher: pick a model and its settings in one window (experimental)
- #1077 feat(arm): build and run Strata on aarch64, including DGX Spark
- #1063 fix(mtp): --batch-mtp exits on the first admission with a draft-vocab head (slot drafters never copy its type)
- #1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer split
- #1048 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installed
- #1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #1039 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #1037 engine: document --spec and --spec-min-p in --help
- #1036 docs: --calibrate loads the model more than once, and starts it when …
- #1033 prefill: gather GPU-resident experts in groups on short prompts too
- #1028 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1024 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)
- #1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)
- #1022 prefill/gdn: the Volta prompt recurrence chain-split (+7-15%) and the opt-in chunked path
- #1021 V100: bit-exact GEMV / MoE operator fusions (experimental build only)
- #1020 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.5% e2e)
- #1017 feature: RDNA3 support / Cache-aware routing
- #1016 Bench: RTX 5060 Ti 16 GB on Ryzen 9 7945HX, IQ3_S
- #1014 vision: downscale large images before CPU encoding (12MP: >8 min → 20 s)
- #1013 serve: a continued batch request reports its own prompt's reuse; input_tokens never 0
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)
- #1001 generate: wait for the residency table's upload before the next window (#871)
- #1000 HIP: gfx1102 (RX 7600) community-validated, with an end-to-end run (follow-up to #192)
- #995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)
- #988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)
- #981 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)
- #970 serve: a tool call written inside the thinking ends it implicitly (#804)
- #955 bench: Arc Pro B65 Gen4 results on patched v0.1.40
- #949 cpu: allow configuring expert-pool tasks per phase
- #946 launcher: pick a model and its settings in one window (experimental)
- #945 setup: the low-RAM and KV-streaming decisions as functions
- #944 hip: gfx11 WMMA prompt attention and selection scorer, gfx1151 hipBLASLt table
- #943 Add experimental deepMoE Vulkan backend for DeepSeek text chat
- #927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854
- #920 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
- #919 hip: Windows support for Radeon gfx1151 APUs (on #895)
- #913 bench: community report, RTX 4060 Ti + Xeon E5-2673 v3 (DDR3, PCIe 3.0 x8), Q2_0 and IQ3_S, 0.1.39 path vs #706 + #764 + a busier tier
- #912 setup: show the GPU's PCIe link in step 1, warn on fewer lanes than the card has
- #911 setup: --inspect, a GGUF's real quantization and whether Strata runs it, from its headers only
- #910 Two-GPU decode: bounded attention merge (bitwise, opt-in) and pipelined PLE read-ahead / early chain (stacked on #905)
- #908 setup: ModelScope as a download source (--source), used when Hugging Face's files do not answer
- #907 calibrate: measure the adaptive expert tier (--adapt-every / --adapt-swaps / --adapt-decay)
- #905 Pipelined windows + async adaptive tier together (stacked on #859 and #876)
- #901 Add ARM64 NVIDIA GB10 build and native-pack support
- #897 setup: the disk check counts only what step 6 still writes
- #895 hip: support Radeon gfx1151 APUs
- #878 # Support sm_60, sm_70, and sm_86 with CUDA 12.6 or 12.8
- #876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)
- #869 serve: opt-in recovery from repeated reasoning passages
- #868 docs: Arc Pro B60 rows and notes, and the B60's PCI id (e211) in setup_intel.py
- #866 sycl: keep cudaStreamQuery's answer in the per-layer ring waits
- #865 ngram: add opt-in Q8_0 PLE table support
- #861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)
- #860 Unsloth UD-Q6_K_XL: Q8_0 PLE rows and Q6_K gate/up experts
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
- #854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)
- #848 Resident RAM mode with a layer split, --vram-reserve-later-mib (#642)
- #844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)
- #839 Add experimental Linux gfx900 setup with wave64 validation
- #835 HIP gfx103x: the prompt path's 16-bit GEMMs in FP16 in and out (prompt reads ~2x on RX 6900 XT)
- #832 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens
- #829 docs: running Strata behind llama-swap
- #823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)
- #815 bench: community results for an RX 6800 on Windows (0.1.39 against 0.1.38)
- #812 serve: Codex MCP tools (web search) fail as 'unsupported call': map namespace__name calls back to their namespace
- #810 serve: accept json_schema roots that are unions of object schemas
- #800 qsa_prompt_attn: Volta (sm_70) m8n8k4 kernel - keep the hi and lo halves in separate accumulator chains
- #794 split: the auto placement sees each stage's PCIe link
- #793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kit
- #788 serve: tolerate broken stdin pipes during engine cleanup
- #785 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installed
- #773 file tier: unbuffered reads on Linux too (O_DIRECT + the kernel's asynchronous reads)
- #766 hip: a gfx1100 hipBLASLt tuning table for 100401 (packaged ROCm 10.0.0)
- #759 Add native Responses API support for Codex clients
- #757 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
- #755 hip: gfx1100 hipBLASLt 1.5.0 tuning table (ROCm 10.2 nightly) — closes the gfx1100 100500 gap
- #750 docs: describe a targeted ROCr USERPTR reclaim workaround
- #734 Reuse long chat history when the last user message is edited
- #726 Add opt-in live cache resizing and desktop resource presets
- #724 Add local PDF, Word and Excel attachments with screenshot previews
- #711 kv: --kv k8v4 streams with --kv-resident
- #705 Feat/k8v4 kv streaming
- #693 prefill: auto:16384 tries every 1024 tokens above 8192, and equal chunks - prompts 21-38% faster from 20K on a 16 GB card
- #686 Add a tiny zero-dependency terminal launcher for Strata
- #685 Add optional Windows HIP vision encoder builds and package validation
- #683 setup: a re-run of setup carries over the hand-edited config blocks (#629)
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #674 bench: community report, RTX 3090 Ti + 2x Xeon E5-2699 v3 (HP Z840), IQ3_S at 262K
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #666 setup, serve: the image encoder's device as its own role (--vision-device)
- #664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)
- #662 setup: prefer uv for the venv and package installs when uv is installed
- #655 sm_75: run the prompt path's BF16 products on the FP16 tensor cores (+15-18 % prefill on RTX 2080 Ti)
- #645 Update the engine from the web app, cautiously
- #636 native: add opt-in GBNF constraints to existing generation
- #634 strata_pack.py build: refuse an output directory that already holds a pack, naming the file
- #630 serve: add opt-in stateless Responses API
- #627 opt for sm_70(just test V100-16g*1 or 2)
- #618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokens
- #615 fix(serve): correct cache usage and timings across reasoning continuations
- #608 setup: a 200K context between the 128K rule and 256K (#406)
- #603 qsa_select: a 1,024-thread top-k past the register kernel's reach (CUDA) - a 243K-token prompt +26% on an RTX 3060
- #598 refill: improve multi-GPU processing speed ~27% PP improvement on 4× RTX 3060
- #595 No-AVX2 CPUs on a native build; run-script context override
- #586 ple: read the PLE table at full BF16 precision (opt-in, alongside FP8)
- #572 serve: return unfinished tool call text as content
- #570 setup: a read-only --gguf-dir no longer stops setup when marking checked shards
- #569 serve, setup: an empty API key is refused from the config file and from setup's --api-key (#213)
- #568 parity tests: ple_parity runs without the missing fixtures, gr_parity and ple_parity wait for their setup
- #567 serve: tokenise the prompt from the last shared prefix, not from the start
- #565 tools/hip: add gfx1100 hipBLASLt tuning table for hipBLASLt 1.2.2 (100202)
- #563 Give the expert cache's VRAM back without unloading the model (#533)
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
- #512 prefill: use active QSA top-k bounds on Turing with large KV capacity
- #508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35
- #504 web: saved chats in a sidebar, kept in localStorage (#361)
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
- #491 docker: GGUF_DIR, RESIDENT_BUDGET_GIB and KV_STREAMING env vars; document the stop timeout
- #472 tests: test_prebuilt_hip_zip passes on Linux (EXE as on Windows)
- #470 setup: a download stopped before its rename is finished without a request (no 5-minute 416 retry loop)
- #469 bench: community results, RTX 2080 Ti 11 GB + Threadripper 3960X, IQ3_S at 262K with KV streaming
- #464 ple: read the n-gram table at full BF16 precision
- #463 decode: wait for an adaptive expert swap before reading the residency table
- #455 setup: mark the Unsloth family NVIDIA only so far in the AMD menu, and point a mixed PC to its NVIDIA card (#429)
- #453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-resident
- #452 qsa_prompt_attn: Q4_0 KV on tensor cores (mode 4) - long prompts with --kv q4_0 ~19% faster
- #450 calibrate: a gpu list is not a layer split (#447)
- #439 prefill: a group's expert gathers in one launch (bit-identical, long prompts +4%)
- #436 serve: a non-streaming request stops when its client disconnects (#430, part of #431)
- #435 setup: resuming a download asks only for the missing space, and a finished .part is not re-requested (#425)
- #434 Add optional unified Strata Manager
- #433 bench: community report, RTX 5090 + Ryzen 9 5950X (AVX2), UD-Q4_K_XL / IQ3_S / Swift IQ3_XXS on engines 0.1.31-0.1.39
- #428 tests: test_setup_golden passes on Linux (<EXE> only as a whole path part)
- #427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)
- #426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)
- #424 fix(setup): drop an engine archive that fails to unpack
- #422 hip: enable experimental Strix Halo gfx1151 source builds
- #417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31
- #413 prefill: DeltaNet recurrence with the three value heads of a key head in one thread (bitwise; 1.4x on a 4080 SUPER, 1.3x on a 3090)
- #409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory
- #407 Adaptive tier: --adapt-decay, and --adapt-tuned (opt-in): 31% fewer misses, -19% CPU pool, -22% PCIe on UD-Q4_K_XL
- #399 fix(setup): drop a refused engine archive instead of reusing it
- #398 setup: stop on EOF instead of accepting prompt defaults
- #395 Feature/nvidia p40
- #389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #386 tools/hip: gfx1201 hipBLASLt table for ROCm 7.2.4 (hipBLASLt 1.2.2), ~1.8x prompt speed on the R9700
- #381 Add HIP image encoder support and Linux AMD GPU telemetry
- #380 HIP: count the desktop's VRAM on Windows (WDDM budget, STRATA_WDDM_BUDGET)
- #378 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #376 HIP: a failed first configure no longer leaves the kernels at -O0
- #374 Prompt path: the first chunk's n-gram (PLE) rows read beside layer 0, 256 at a time (same output)
- #372 Prompt path: an MMQ group's experts gathered in one launch, after one wait, released by one event (-9% on a 32K prompt, same output)
- #366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Next
- #362 File tier: read the GGUF in place unbuffered when the file cache cannot keep it beside the RAM budget (Windows; #286 ported into the existing tier)
- #358 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #357 Start: read the expert arena unbuffered when the file cache cannot help (Windows; part 1 of #285)
- #356 HIP: build the AMD engine on Windows
- #346 setup: don't list the shared Chat settings file as a model config (KeyError: 'exe')
- #339 hip: hipBLASLt tuning table for gfx1201 (R9700) with ROCm 10.2.0a nightly (hipBLASLt 1.5.0)
- #334 serve: enforce structured JSON in the native sampler
- #330 V100 (sm_70): the compute-capability floor drops from 7.5 to 7.0
- #325 AMD working on Windows (9070XT tested)
- #324 fix(setup): install the engine, model files and packages the checkout was tested with
- #323 hip: add experimental RX 5500 XT 8 GB (gfx1012) support
- #314 feat: Qwen3.8-Flash-Next dense model, native Qwen3.5 attention kernel, strata-dense CLI
- #313 RDNA3 WMMA kernels for the gfx1100 prefill (opt-in, runtime-gated on gfx11; rebased on 0.1.30)
- #311 HIP: support RDNA2 gfx1030 working and gfx103X is untested.
- #296 Add OrcaRouter Q4_K_S support (0.1.30)
- #295 Pascal build runs (STRATA_EXPERIMENTAL_SM60)
- #292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rows
- #289 Start: name the GPU and stop at once when the build has no code for it; STRATA_EMULATE_CC
- #287 setup: --draft-vocab cyrillic (English/code + the whole Cyrillic script)
- #286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.bin
- #285 Start: read the expert arena unbuffered (the GGUF too, fixes #230), register it while it loads, on its own thread
- #279 Windows: leave 1500 MiB of VRAM unused by default (decode stalled with 700)
- #275 Add bounded disk persistence for evicted conversations
- #264 tests: make iq_parity reproducible from a fresh checkout
- #262 HIP: gfx1201 (Radeon AI PRO R9700) support, faster packed byte intrinsics, and a gfx1201 hipBLASLt table
- #256 HIP: support RDNA4 gfx1200 (RX 9060 XT)
- #247 HIP: Windows and gfx1201 (RX 9070)
- #242 IQ kernels: decode each weight part once for every column and entry, bitwise identical
- #241 Faster grouped and per-hit Q2_0 expert kernels, bitwise identical
- #240 gemm: name the cuBLAS/CUDA call that failed at init, and flag resource shortages
- #229 multi-GPU: second GPU as an opt-in expert-cache tier with P2P rows (--peer-device), part 1 of #204
- #227 AVX1: the Q2_0 expert kernels for CPUs that have AVX but no AVX2
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)
- #202 ple: --ple-io ram locks the n-gram table in RAM (no SSD read on the prompt/token path)
- #194 serve: clear the error string at the start of every request (fixes #183)
- #192 HIP: support RDNA3 gfx1101/gfx1102 (RX 7600/7700/7800), not only gfx1100
- #190 Spill evicted conversation snapshots to a bounded optional disk cache
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)
- #175 Preserve independent conversation caches across interleaved requests
- #166 build: replace CUDART_INF_F with INFINITY for HIP-only toolchains (fixes #157)
- #158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUs
- #151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setup
- #140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S support
- #136 Linux prebuilt from CI: CUDA 12.8, sm_80/86/89 + PTX, attached to each release
- #134 setup: count the 3-bit models' 262K context against RAM, not a fixed 90 GB
- #130 setup: Volta (V100, sm_70) and Turing as an experimental build path
- #126 Windows: the expert load runs 24x slower from Task Scheduler (0.05 vs 1.42 GiB/s); a hint names the cause
- #120 `--kv k8v4`: hybrid KV cache — INT8 K + Hadamard-rotated Q4_0 V (816 B/cell)
- #111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU
- #109 decode: batch the verify window's per-token kernels (bit-identical, +~10% decode)
- #108 prefill: bit-identical kernel speed-ups (batched indexer append, parallel GDN conv, column-split GDN recurrence, one-launch embedding gather)
- #105 The Monitor's GPU readings on ROCm: an amdgpu sysfs backend, and the disk rate without psutil
- #96 Add docker build
- #94 Add experimental gfx1100 HIP backend for RX 7900 XTX
- #92 setup: name a missing or short shard with its numbers; verify reads back the manifest sha256
- #91 GGUF reader: refuse a duplicate tensor name at open, naming the tensor and the file
- #88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)
- #87 Turing port: run on RTX 20 (sm_75) and newer
- #84 EXPERIMENTAL Rope scaling: contexts past the trained 262K (none / linear / YaRN)
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)
- #67 Support OrcaRouter IQ3_XXS with explicit BF16 compatibility packing
- #59 Fix/53 sampler penalties. sampler: penalties apply once, as in llama.cpp's default chain (#53); guard the unsized penalty bitmap
- #50 tools: add mmproj quantization utility
- #44 generate: probe the real host->device bandwidth for pcie_frac
- #24 serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting
- #16 Add optional multi-GPU expert caches and compact transfers
- #12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…