标签 / Windows
Windows
Auto topic: Windows
Issues
- #1495 [Feature]: experimental per-stage KV-grow for a two-GPU layer split at native 262K
- #1487 Windows/CUDA 11.8 Volta integration: focused patches and reproducible handoff
- #1474 hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
- #1468 prefill copy_i32: illegal memory access on long prompts when the prefill chunk is large (auto:32768)
- #1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combine
- #1442 Feature request: support chat title generating from Xcode's agent
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
- #1413 `--batch-groups G>1` + `--batch-mtp` is silently inert: the pipelined path never calls `batch_step`, but still builds per-slot drafters
- #1412 Windows: `SeLockMemoryPrivilege` moves the expert arena to 2 MiB large pages — but the token is issued at logon (1314 ≠ 1450)
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying
- #1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request
- #1397 SYCL port on Windows/OpenCL: 0.1.40-sycl decodes at 22.3 tok/s with --spec 2 where 0.1.39-sycl does 53.6 (and short runs vary 40%)
- #1393 Support for AMD Radeon 890M APU
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by default
- #1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it
- #1369 Two 32K prompts on one RTX 5090: overlapping them matched back-to-back wall time
- #1366 [AMD/gfx1031/Windows] Self-built engine crashes with an access violation inside amdhip64_7.dll before any output
- #1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40
- #1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stack
- #1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)
- #1344 Server: Multiple API key(llamacpp format) will seen as one long key
- #1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the model
- #1317 serve: an untimed read of the engine/vision pipe can hold the request turn forever (intermittent hang; the #481 lost-step class)
- #1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)
- #1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
- #1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
- #1299 AMD vison encoder on Windows
- #1286 PCIe Issue with Strata and RTX 3060
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1277 HIP: fused IQ expert kernels via RDNA4 int8 WMMA (RX 9060 XT / gfx1200)
- #1276 UPDATE.bat cannot update existing checkouts after the main history rewrite
- #1275 Windows HIP: --prefill auto can pick a chunk whose VRAM borrow leaves too little for the verify captures — silent engine exit at serve init
- #1272 0.1.40: with the shared-expert stream fork off by default, decode on 2x RX 6900 XT (gfx1030, Linux) is 5-13% slower; STRATA_SH_STREAM=1 restores it, and the #884 stall does not come back with it
- #1261 461
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
- #1252 DraftPolicy: a lookup window size priced during a slow stretch stays over-priced for the rest of the process
- #1239 v0.1.40 two-GPU tester results on a P40 + RTX 3070: 150-request soak clean, --pipeline-windows 2 +9% / +15% decode in 5 of 5 pairs
- #1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32
- #1234 Windows/RTX 4090 (0.1.39): stall report stage "decode -1" for 61 s; 0 layers served - thread stacks attached
- #1233 ``` Windows + Strix Halo (gfx1151, Radeon 8060S): UD-IQ4_XS runs end-to-end — decode flat 29.5-30.4 tok/s at 1K-32K, but a GPU driver timeout (TDR) during a cold 32K prefill ```
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1208 HIP layer split / gfx1030: batched prompt path produces NaN logits (output "!!!!") unless STRATA_HIP_PROMPT_F16=1 and the gfx1030 card is the first stage
- #1200 RTX 5060 Ti 16 GB / Windows / IQ3_S: 36–42 tok/s — tuning directions & method (local gate scan, large pages, PLE FP8)
- #1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)
- #1177 Feature request: add timestamps to engine log lines
- #1165 Ship an MMQ-enabled engine build (or opt-in loading) for k-quant prompts: a full-RAM machine is GPU-bound on the FP16 dequant path
- #1155 [Windows / AMD] --vision cpu is still unreachable: hip_vision()'s WIN gate, find_vcvars()'s VS-18 range, and no encoder in the ready-made engine
- #1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
- #1138 Windows: higher process/thread priority improves decode throughput under CPU contention, with a measured background-work trade-off
- #1128 0.1.40 --kv-grow: a short request trims the K/V to 8192 cells and drops the conversation cache, so the next 128K+ continuation is re-read in full (20.9% of continuations, RTX 5090, IQ3_S)
- #1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML type
- #1118 Batch slots: the engine stalls ("no progress for 60 s … reading the prompt (batched)") when a long prompt is read while another slot decodes; STRATA_BATCH_DECODE_SHARE=0 avoids it
- #1103 hip: gfx1030 (RX 6800 XT) intermittent verify timeouts (#267): ROCm 7.14 gfx103X wheels regression, fixed by 7.13
- #1085 Performance regression after updating to v0.1.40 - RTX 5090
- #1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40
- #1072 Vision: >64 unique images can evict .sve files still referenced by the current request
- #1070 New Feature - add option to change location of model
- #1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)
- #1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)
- #1005 CRLF line-break tokens (`\r\n`, ids 317/845/8301) get far too little probability; everything else matches llama.cpp
- #990 MCP install planning rejects supported AMD CPU vision (and retains the old Windows HIP restriction)
- #985 Also allow Visual Studio 2026 to build from source
- #980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missing
- #977 --check prints [ok] for below-floor RAM and a misleading 'This PC can run Strata' verdict
- #975 setup.get_prebuilt_hip() zip-unpack branch returns None on Linux (tools/test_setup_amd.py)
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
- #961 Windows (WDDM): verify window wedges nvlddmkm — 0x141 engine timeout → 0x116 TDR failure (driver defect, not a Strata bug)
- #953 Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decode
- #952 BrokePipeError when terminating Strata on Linux
- #925 Setup: show progress in the browser from the first second (status page with steps, speed and ETA)
- #922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runs
- #918 Windows: Radeon 8060S (gfx1151) builds, selftests and passes the HIP ctest on top of #895
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
- #915 Windows AMD report: RX 6800M (gfx1031, laptop) works with a self-built HIP engine — please add 0x73DF / gfx1031 to the Windows path
- #906 Older PC (Haswell, DDR3, PCIe 3.0 x8): +15% decode on Q2_0 and IQ3_S from #706, #764 and a busier adaptive tier
- #881 [Windows / AMD] RX 6800M (gfx1031) runs with a self-built HIP engine — plus 2 Windows bugs found (packaging encoding, vision `vcvars`)
- #872 Session RESTARTS COMPUTER after 25-30 mins (gfx1100)
- #871 Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)
- #857 "parallel": 2 lowers total throughput on a single 48 GB card (100% expert-cache hits): keep MTP drafts in batch windows?
- #840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaningless
- #836 Milestone: Add Apple Silicon Cross-Platform Support
- #830 Let a one-shot request skip the prompt-cache work (~45 ms per call that is never reused)
- #825 Results for --model UD-IQ4_XS
- #821 ROCm on Windows ZIP
- #816 0.1.39: decode on an RX 6800 (gfx1030, Windows) is 12-22% slower than 0.1.38; STRATA_SH_STREAM=0 removes it
- #806 Strange behavior for x2 RTX 5060 Ti (16Gb each) and UD-Q4_K_XL
- #804 Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`
- #803 Perplexity ~7-10% above llama.cpp on the same GGUF (Qwen3.8-Flash-Next), also at 2K context, every version tested
- #798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`
- #795 Main engine strata.exe crashes with STATUS_ILLEGAL_INSTRUCTION (0xC000001D) on non-AVX-512 CPU - 20 WER crashes across v0.1.35/v0.1.38, fault offsets cluster in fixed code regions
- #781 1M context: identical expert work per window, but the VRAM-touching stages grow 7-8x
- #779 laptop crash with latest
- #776 --batch on a 100%-resident layer split: engine dies at "captured the batch window over slots 0,1" (zero-doorbell path?)
- #775 Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile
- #756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)
- #735 can't load vision encoder
- #730 WARNING: the RAM budget (--resident-budget-gib) cannot be kept - experimental Q4
- #728 Infinite loop issue
- #710 v0.1.38: cross-request tool loop repeats an ineffective repair 98 times despite a thinking budget
- #697 Windows HIP / gfx1201: prompt stalls at `GPU rang 0`; post-store `__threadfence_system()` resolves it
- #691 Windows: generation slows when the server window is minimized (EcoQoS moves threads to E-cores) — fix included
- #675 Feature request: `logprobs` / `top_logprobs` on `/v1/chat/completions` (typed decisions from one forward pass)
- #671 serve: ordered lists whose items are separated by blank lines render every item as "1."
- #670 Keeping the old engine when an update succeeds but turns out bad - and a question about where the web app should call it
- #669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
- #658 Proposal: single-GPU prefill buffer planning and RAM budgeting for conversation caching (RTX 5090 measurements)
- #654 HIP (gfx1200) RuntimeError Exception Code: 0xC0000005
- #642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)
- #633 The pinned expert arena is allocated without a host-RAM check: a container limit or a low MemAvailable ends in the OOM killer mid-load, not a message (the file tier already has the probe)
- #629 Re-running setup with a different --context silently discards mcp_servers, mcp and sampling from the run config
- #620 IQ3_S: "native head upload: out of memory" at startup when the desktop runs on a second GPU (more free VRAM on the card)
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #606 After a 36,689-token request at 155k, every later request answers one repeated token until the engine reloads
- #601 Linux build fails on Ubuntu 26.04 (glibc 2.43) with CUDA 12.9 (nvcc host-math conflict)
- #597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090
- #592 serve: a malformed "tools" value kills the request thread (connection reset / 502 instead of a 400)
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)
- #577 0.1.38: UD-Q4_K_XL prompts 15-40% slower on a 96 GB PC, from the unbuffered file tier (STRATA_UNBUFFERED_LOAD=0 restores them)
- #564 Gui settings
- #556 EXL3 (exllamav3 trellis) weight support
- #551 Wrong UI on Windows?
- #548 decode_cluster_parity: graph replays race their input upload (pageable cudaMemcpy + non-blocking stream)
- #535 Thinking speed suddenly drops to 0.5 token/s
- #533 Hot VRAM resize: let the engine yield VRAM to other apps and take it back
- #530 high/xhigh reasoning_effort can silently run to max_tokens with empty content when reasoning_budget_tokens isn't set
- #528 Conversation cache: decode drops ~4-6x on continued conversations (0.1.36, Windows, RTX 5090, IQ3_XXS)
- #519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt path
- #515 Feature request: Intel Arc B390 / Panther Lake iGPU support on Windows (64 GB shared RAM)
- #511 Engine stopped unexpectedly
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
Pull requests
- #1499 core: write reduced helper rows directly to mapped output
- #1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegister
- #1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)
- #1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMA
- #1493 Experimental cooperative memory relief and resumable text requests
- #1492 fix(serve): save and restore sessions without an MTP drafter
- #1481 web: add Copy below code blocks (#1257)
- #1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)
- #1477 docs: correct pipeline async compatibility (#1447)
- #1476 Update GLM port to Project Maya v1.0.4
- #1471 feat: optionally offload idle session state while keeping model weights loaded
- #1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3
- #1465 cuda: retain HC norm products in existing scratch
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K
- #1458 helper cache: configurable VRAM allowance with allocation-floor checks
- #1457 prefill: reuse shared scratch for the HC read
- #1451 prefill: allocate PLE host buffers only when needed
- #1450 web: wide layout — a four-column single page on very wide screens, and a card look for the chat and the drawer
- #1434 telemetry: GPU load, VRAM, temperature and power for AMD cards on Windows (#1380)
- #1426 dflash: experimental standalone DeepSpec DFlash block drafter for Qwen3.8-Flash-Next (greedy, correctness-first)
- #1424 Pascal GP100: a Q8_0 decode GEMV (STRATA_Q8_SM60=1, opt-in) and IQ3_XXS / IQ3_S tables in shared memory
- #1420 Add optional GLM-5.3-Flash support from Project Maya
- #1419 bench: replay token IDs and verified windows in serve decode
- #1418 cuda: opt-in MMVQ activation reuse across output rows
- #1417 prefill: scope each split-stage loan estimate to its device
- #1416 prefill: the CPU share's activations quantized on the GPU, copied down beside the routing sync (builds on #1414)
- #1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer split
- #1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweep
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1399 qsa_select_bench: accuracy floor per sample (fails on an RTX 5090 / MSVC build)
- #1395 hip: a hipBLASLt tuning table for gfx1150 (Radeon 890M), and gfx1151's exact speed switches on gfx1150
- #1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
- #1384 test: fix narrowing byte fixtures in Windows conversation test
- #1383 generate: say that CUDA_LAUNCH_BLOCKING=1 hangs the verify windows (#1341 #964)
- #1379 prefill: STRATA_PREFILL_CPU_SHARE=auto times layers with and without the share and shares only while that is faster (#1282)
- #1377 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.2
- #1376 bench: add RX 7900 XT results for IQ3_S and IQ3_XXS
- #1374 experts: IQ1_S on GPU/AVX-2, IQ1_S+IQ1_M on AVX-512, and a measured PCIe share
- #1370 prefill: opt-in STRATA_HC_UPMIX=1 on CUDA - the hyper-connection up projection with gr_mix_r as its epilogue (sm_80+)
- #1362 serve: preserve Markdown state across skipped tool-call newlines
- #1361 tools/hip: gfx1151 hipBLASLt 100500 tuning table (Windows ready-made …
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)
- #1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4
- #1354 docs: the expert cache is a budget, and on Windows an over-sized one pages instead of failing
- #1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #1343 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers
- #1332 calibrate: the PCIe share sweep reaches 1.0, where no expert goes to the CPU pool
- #1331 Aporte/disk mirror y sesiones agnosticas
- #1329 Fix batch MTP shared draft head type
- #1328 serve + engine: constrained decoding for JSON formats; accept tools together with json_schema
- #1327 serve: opt-in finite video input
- #1325 file tier (Windows): read each expert from experts.bin and a mirror on a second drive at once
- #1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)
- #1323 prefill: batched unbuffered stager reads; close the experts.bin view once reads are unbuffered (port of #833)
- #1319 serve: add opt-in bounded Responses history and continuation
- #1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #1295 Experimental native Windows path for Intel Arc
- #1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context
- #1287 feat(vision): support Windows HIP image encoding
- #1282 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
- #1263 Community benchmark: RTX 5090 Laptop GPU (24 GB), Windows 11, Strata 0.1.40.1
- #1262 STRATA_HC_REQ8: hyper-connection projections requantized to int8 + fp32 scale per 32 at load (rebased #704)
- #1260 tests: order MMQ parity transfers on the compute stream
- #1253 Enable serial multi-GPU batch MTP
- #1249 Pipelined batch: two requests run at half speed (pad rows in group windows, unbalanced slot choice)
- #1243 Bench/2026 10 05 community rtx 5080
- #1241 serve: the Responses API reads Codex's additional_tools input items as tools (#782)
- #1232 enable --kv-grow with --batch
- #1231 kv-grow: a run that cannot lend cache slots maps the whole window up front instead of writing past its first 16K cells
- #1228 serve: list items separated by blank lines are one list, and a list continued later keeps its numbers (#671)
- #1227 mcp: the install plan follows setup's AMD rules - Windows AMD and vision=cpu on Linux are planned (#990)
- #1224 serve: the Monitor's AMD GPU readings on Windows through ADLX (load, VRAM, temperature, power)
- #1220 fix #1121: preserve shared draft-head metadata in batch MTP slots
- #1218 Check the ready-made engine against GitHub's SHA-256 before installing it
- #1216 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens (reland)
- #1207 fix(serve): keep cached images alive through request preparation
- #1205 update.sh: a rewritten history is not the user's fault, and the messa…
- #1204 serve: a Chinese / English language switch for the web app
- #1199 Community benchmark: RTX 3090 eGPU + 64GB, IQ2_XS vs IQ3_XXS
- #1185 verify: a window graph that finds no VRAM frees older batch slot layouts' graphs instead of failing (#997)
- #1184 AMD HIP: the primary ordinal, and the helper caches beside the resident RAM mode
- #1181 perf(expert): the verify window's GPU plan in O(n) instead of O(n^2)
- #1179 batch: with --batch-groups, a group's lowest free slot first (+14-22% at 2-3 clients)
- #1173 bench: community report, RX 7900 XTX on Windows 11, IQ3_S, engine 0.1.40
- #1170 test(gr): check the fused multi-token read at T=4 and T=2
- #1166 cpu pool: --host-core sibling, the host off the interrupts' logical processor (hybrid CPUs too)
- #1164 conversation cache: a shared prefix is copied out of a parked conversation instead of taking it over, so the parked one keeps its history
- #1163 conversation cache: a re-parked conversation reserves at most 1/8 of each K/V buffer instead of up to 16 MiB, so fewer get evicted
- #1162 Verify the downloaded engine against GitHub's SHA-256 (replaces #645)
- #1156 setup: reach the CPU image encoder on Windows + AMD (#1155, #881)
- #1154 decode: support a three-GPU two-window pipeline on 3 P4s
- #1152 Fix Windows AMD telemetry and add optional Arabic UI
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
- #1136 Add opt-in durable chat history, legacy import and browser compaction
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
- #1130 Windows: apply commit limits after allocation failure
- #1127 hip: rebased support for Radeon gfx1151 APUs
- #1125 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)
- #1122 Pipelined windows + async adaptive tier together (--pipeline-windows 2 --adapt-async 1)
- #1120 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, opt-in)
- #1119 Add local PDF, Word and Excel attachments with screenshot previews
- #1117 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in; resubmission of #1030 on the new main)
- #1115 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #1114 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1110 fix(mtp): --batch-mtp exits on the first admission with a draft-vocab head (slot drafters never copy its type)
- #1107 prefill: gather experts in groups on short prompts too
- #1102 fix: batch host paths honor refreshed residency (deadlocks the engine during a prompt loan)
- #1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)
- #1092 launcher: pick a model and its settings in one window (experimental)
- #1090 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #1089 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing traces
- #1086 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
- #1078 bench: RX 6800M (gfx1031) community report - self-built HIP engine on Windows
- #1064 Windows: apply commit limits after allocation failure
- #1063 fix(mtp): --batch-mtp exits on the first admission with a draft-vocab head (slot drafters never copy its type)
- #1060 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)
- #1048 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installed
- #1044 serve: image sources - network paths refused, URLs capped and fetched outside the FIFO, no local files for other origins' pages
- #1041 serve: an image in an Anthropic tool_result reaches the encoder (Claude Code's Read of a picture)
- #1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #1039 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #1033 prefill: gather GPU-resident experts in groups on short prompts too
- #1032 serve: add opt-in bounded Responses history and continuation
- #1030 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in, CUDA VMM)
- #1025 fix: unblock all-resident batch windows (batch verify waited for a doorbell the graph never publishes)
- #1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)
- #1011 kv: one pinned KV pool shared by the sessions (--kv-pool-tokens; 4 lanes x 256K pin 6.24 GiB instead of 15.5)
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)
- #1003 resident RAM mode: release each layer's mapped experts as it is copied (Windows)
- #1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)
- #1001 generate: wait for the residency table's upload before the next window (#871)
- #998 serve: an engine that exits after an unread ERR line says why (#997 #890)
- #993 bench: community report, RTX 5070 Ti + Ryzen 9 5900X, 64 GB DDR4-3600 (IQ2_XS / IQ3_XXS / IQ3_S, six-arm A/B)
- #991 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing traces
- #987 serve: /props reports chat_template_caps so Zed offers tools (#986)
- #960 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #956 serve: an API key with spaces around it or outside Latin-1 works, and a refused key is not called a missing one (#725)
- #949 cpu: allow configuring expert-pool tasks per phase
- #946 launcher: pick a model and its settings in one window (experimental)
- #936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…
- #933 CPU trellis kernels (AVX2 / AVX-512) for TQ2_T / TQK6 / TQK7 expert-cache misses (stacked on #928)
- #919 hip: Windows support for Radeon gfx1151 APUs (on #895)
- #910 Two-GPU decode: bounded attention merge (bitwise, opt-in) and pipelined PLE read-ahead / early chain (stacked on #905)
- #907 calibrate: measure the adaptive expert tier (--adapt-every / --adapt-swaps / --adapt-decay)
- #905 Pipelined windows + async adaptive tier together (stacked on #859 and #876)
- #904 Verify window: PDL on sm_90+, graph branches, batched Q4_0 KV append and PLE post-ops (bit-identical)
- #903 prefill: a prompt reads in equal chunks; a short streaming tail stays (#693, rebased)
- #902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)
- #895 hip: support Radeon gfx1151 APUs
- #889 sycl: read doorbell flags with an uncached L1+L3 hint (the GPU never saw the host's store on an Arc Pro B60)
- #876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)
- #873 Fix first-kernel-launch crash on Windows (ROCm/Adrenalin WDDM startup race)
- #865 ngram: add opt-in Q8_0 PLE table support
- #864 perf: opt-in resident expert exchange buffer rotation
- #861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)
- #859 Layer split: pipelined verify windows for one conversation (--pipeline-windows, opt-in)
- #853 cuda: specialize hyper-connection up for short verify windows
- #851 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
- #848 Resident RAM mode with a layer split, --vram-reserve-later-mib (#642)
- #846 Keep MTP drafting in concurrent batch slots
- #845 verify: batch windows honor the all-resident stage (no host doorbells)
- #844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)
- #839 Add experimental Linux gfx900 setup with wave64 validation
- #837 server: "return_progress": true puts the prompt's PP line on the stream as llama.cpp's prompt_progress
- #833 Prefill: batched unbuffered expert reads, close the experts.bin view; opt-in RAM budget with helper caches
- #832 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens
- #826 verify: the shared-expert stream fork starts off on HIP (fixes the 0.1.38 -> 0.1.39 decode regression, #816)
- #818 Windows: opt-in release of mapped expert pages after GPU uploads
- #815 bench: community results for an RX 6800 on Windows (0.1.39 against 0.1.38)
- #813 rtx6000fix
- #811 Add an RTX 5090 + Ryzen 7 9700X IQ3_S benchmark on engine 0.1.39
- #807 cuda: opt-in batched expert uploads to reduce host submission overhead
- #802 Run IQ1_S (Unsloth UD-IQ1_S) on the GPU and AVX-2, and read the expert file tier faster
- #799 docs: the expert cache size is a budget, and on Windows an over-sized one pages instead of failing
- #796 generate: one startup VRAM plan before the expert cache is committed (#765)
- #793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kit
- #792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layer
- #789 prefill: overlap shared-expert work with routed expert uploads
- #788 serve: tolerate broken stdin pipes during engine cleanup
- #785 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installed
- #783 perf(cuda): fused decode/verify/MTP kernels and graph launch reductions on 0.1.39
- #780 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #777 Add an RTX 5090 + Ryzen 7 9700X IQ3_XXS benchmark on engine 0.1.39
- #773 file tier: unbuffered reads on Linux too (O_DIRECT + the kernel's asynchronous reads)
- #764 decode: STRATA_ADAPT_LAG - #463's reproducibility without its stall
- #759 Add native Responses API support for Codex clients
- #751 Persist conversation KV state in a bounded disk LRU cache
- #750 docs: describe a targeted ROCr USERPTR reclaim workaround
- #749 Respect Windows commit capacity
- #748 serve: invalidate the engine after fatal verification timeouts
- #744 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #733 Pipeline CPU expert work without a batch-wide phase barrier
- #727 Add opt-in durable chat history, legacy import and browser compaction
- #726 Add opt-in live cache resizing and desktop resource presets
- #724 Add local PDF, Word and Excel attachments with screenshot previews
- #719 Add sandboxed HTML previews and scrollable chat code blocks
- #717 fix: make the CUDA language link the shared runtime (Windows sm_70 LNK1169, #585)
- #712 Fix Windows AMD telemetry and add optional Arabic UI
- #711 kv: --kv k8v4 streams with --kv-resident
- #707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)
- #704 STRATA_HC_Q8: hyper-connection projections as int8 + fp32 scale per 32
- #699 Linux: read ahead at startup (cold start ~920 s → 70 s)
- #698 Bench: RTX 5090 / 9950X3D (IQ3_S, three-pass A/B) + RTX 2080 Ti / 9900KF (IQ3_XXS, Turing)
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)
- #693 prefill: auto:16384 tries every 1024 tokens above 8192, and equal chunks - prompts 21-38% faster from 20K on a 16 GB card
- #686 Add a tiny zero-dependency terminal launcher for Strata
- #685 Add optional Windows HIP vision encoder builds and package validation
- #683 setup: a re-run of setup carries over the hand-edited config blocks (#629)
- #682 serve: a Chinese / English language switch for the web app
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)
- #662 setup: prefer uv for the venv and package installs when uv is installed
- #660 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
- #659 Fix WDDM host pinning for gfx1031 native inference
- #656 serve: cooperative prefill preemption at safe chunk boundaries (--prefill-preempt)
- #650 pinned: back the Linux expert arena with transparent huge pages
- #645 Update the engine from the web app, cautiously
- #640 STRATA_ARENA_MMAP=1: the expert arena as a read-only mapped file (Linux), for small-RAM machines
- #636 native: add opt-in GBNF constraints to existing generation
- #630 serve: add opt-in stateless Responses API
- #628 bench: report Coder IQ1_M on Threadripper 3990X and RX 6900 XT
- #627 opt for sm_70(just test V100-16g*1 or 2)
- #626 cpu: preserve full thread affinity on Windows and Linux
- #619 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers
- #618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokens
- #611 cpu: group ARM capacities and honor the pool's reserved host core
- #603 qsa_select: a 1,024-thread top-k past the register kernel's reach (CUDA) - a 243K-token prompt +26% on an RTX 3060
- #594 serve: read the request body an answer left unread, so the close is not a reset
- #586 ple: read the PLE table at full BF16 precision (opt-in, alongside FP8)
- #575 qsa select: split the top-k past the register fit - 9% faster long-prompt prefill
- #567 serve: tokenise the prompt from the last shared prefix, not from the start
- #563 Give the expert cache's VRAM back without unloading the model (#533)
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
- #553 serve: image sources - network paths refused, URLs capped and fetched outside the FIFO, no local files for other origins' pages
- #540 volta (sm_70) and turing (sm_75): faster prompt reading - BF16 projections on the FP16 tensor cores, and a faster pre-Turing attention kernel
- #536 decode_cluster_parity: the graph replays wait for their uploads (the test failed now and then)
- #532 tests: prefill_mmq_kquant_test waits for its uploads before the kernels
- #531 multi-GPU: second GPU as an opt-in expert-cache tier (--peer-device), resubmit of #229 on v0.1.36
- #529 serve: an image in an Anthropic tool_result reaches the encoder (Claude Code's Read of a picture)
- #527 Experimental P100: pinned RAM complement for fixed layer-split CPU/GPU inference
- #512 prefill: use active QSA top-k bounds on Turing with large KV capacity
- #504 web: saved chats in a sidebar, kept in localStorage (#361)
- #500 perf(cpu-pool): eliminate host serialization barrier for intermediate activation quantization with adaptive threshold
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
- #488 pinned: honor STRATA_NO_LARGEPAGES on Linux; name the hugetlb pool shortfall
- #484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)
- #483 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #482 Networked Pool support
- #480 serve, web: a space-free vision temp dir, lazy vision, and a Chinese UI
- #472 tests: test_prebuilt_hip_zip passes on Linux (EXE as on Windows)
- #470 setup: a download stopped before its rename is finished without a request (no 5-minute 416 retry loop)
- #462 diagnostics: dump every verify window's logits, so greedy decode can be bisected
- #456 Monitor: the conversation cache card (parked conversations from the engine's CACHE lines)
- #453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-resident
- #452 qsa_prompt_attn: Q4_0 KV on tensor cores (mode 4) - long prompts with --kv q4_0 ~19% faster
- #443 kernels: add opt-in small-window GR and one-warp MMVQ experiments
- #442 hip: add community gfx1012 support with version-gated legacy compatibility
- #441 Contrib/non mtp serving
- #440 bench: community report, RTX 5090 + Ryzen 9 9950X3D, IQ3_S at the full 262K window; prefill auto:32768 A/B; conversation-cache 32 GiB demo
- #439 prefill: a group's expert gathers in one launch (bit-identical, long prompts +4%)
- #436 serve: a non-streaming request stops when its client disconnects (#430, part of #431)
- #433 bench: community report, RTX 5090 + Ryzen 9 5950X (AVX2), UD-Q4_K_XL / IQ3_S / Swift IQ3_XXS on engines 0.1.31-0.1.39
- #428 tests: test_setup_golden passes on Linux (<EXE> only as a whole path part)
- #427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)
- #426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)
- #424 fix(setup): drop an engine archive that fails to unpack
- #416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budget
- #415 IQ4_XS expert rows on the AVX-2 path: the multi-token kernel this format never had
- #413 prefill: DeltaNet recurrence with the three value heads of a key head in one thread (bitwise; 1.4x on a 4080 SUPER, 1.3x on a 3090)
- #407 Adaptive tier: --adapt-decay, and --adapt-tuned (opt-in): 31% fewer misses, -19% CPU pool, -22% PCIe on UD-Q4_K_XL
- #399 fix(setup): drop a refused engine archive instead of reusing it
- #398 setup: stop on EOF instead of accepting prompt defaults
- #381 Add HIP image encoder support and Linux AMD GPU telemetry
- #380 HIP: count the desktop's VRAM on Windows (WDDM budget, STRATA_WDDM_BUDGET)
- #378 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #377 HIP: time the PCIe probe on the host on Windows
- #376 HIP: a failed first configure no longer leaves the kernels at -O0
- #374 Prompt path: the first chunk's n-gram (PLE) rows read beside layer 0, 256 at a time (same output)
- #373 HIP: build with Visual Studio 2026's C++ library (MSVC 14.51+)
- #372 Prompt path: an MMQ group's experts gathered in one launch, after one wait, released by one event (-9% on a 32K prompt, same output)
- #363 experts: smaller and fewer launches for the verify window's PCIe call
- #362 File tier: read the GGUF in place unbuffered when the file cache cannot keep it beside the RAM budget (Windows; #286 ported into the existing tier)
- #358 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #357 Start: read the expert arena unbuffered when the file cache cannot help (Windows; part 1 of #285)
- #356 HIP: build the AMD engine on Windows
- #353 NVFP4 routed experts on 0.1.35 (experimental): 120a only for the W4A4 unit (rebase of #292)
- #346 setup: don't list the shared Chat settings file as a model config (KeyError: 'exe')
- #334 serve: enforce structured JSON in the native sampler
- #333 serve: add lazy startup and reliable Windows model cleanup
- #332 serve: add a standalone API request monitor
- #325 AMD working on Windows (9070XT tested)
- #324 fix(setup): install the engine, model files and packages the checkout was tested with
- #323 hip: add experimental RX 5500 XT 8 GB (gfx1012) support
- #317 ple: keep the SSD awake while rows are read (STRATA_SSD_KEEPALIVE)
- #315 hc: the multi-token hyper-connection read split finer
- #314 feat: Qwen3.8-Flash-Next dense model, native Qwen3.5 attention kernel, strata-dense CLI
- #302 HIP/Windows: batch expert copy dependencies in long prefill
- #296 Add OrcaRouter Q4_K_S support (0.1.30)
- #294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windows
- #293 KV: optional Hadamard rotation for int8 K/V (STRATA_KV_ROT=1), with the drafter and batched verify made rotation-consistent
- #292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rows
- #291 PLE: the n-gram table in FP8 E4M3, as Qwen ships it (opt-in)
- #290 Token embedding in BF16 from the checkpoint (--embd-gguf)
- #286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.bin
- #285 Start: read the expert arena unbuffered (the GGUF too, fixes #230), register it while it loads, on its own thread
- #284 Verify: the commit graph no longer waits; it overlaps the MTP draft
- #283 Prompt path: BF16 projections get the activation's BF16 remainder too
- #282 prefill auto: chunks up to 32768 (was 8192), +15% on a 32K prompt
- #281 numerics: saturate the prompt path's FP16 SwiGLU products; log1pf in the unfused GDN gate
- #280 RoPE: every kernel reads the session's float64 angle table
- #279 Windows: leave 1500 MiB of VRAM unused by default (decode stalled with 700)
- #278 serve: count_tokens, Anthropic thinking only when asked, STRATA_REQUEST_LINES, a relative exe
- #275 Add bounded disk persistence for evicted conversations
- #272 feat(cpu): add hybrid architecture awareness and --pool-affinity for P/E-core CPUs
- #269 Split prefill: run the whole prompt on the main GPU (--prefill-main)
- #255 Q8 0 support
- #247 HIP: Windows and gfx1201 (RX 9070)
- #242 IQ kernels: decode each weight part once for every column and entry, bitwise identical
- #241 Faster grouped and per-hit Q2_0 expert kernels, bitwise identical
- #239 verify: stage the mode-2 PCIe share on the copy stream inside the captured window
- #235 verify: STRATA_LOGPOS hook for per-position log-probabilities (items 12-13 of #84)
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)
- #216 Multi-GPU: the session carve and the per-stage prompt loans
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)
- #207 E-2 on the AVX-2 path too: prefetch the i-quant expert rows (same switch, CPUs without AVX-512)
- #190 Spill evicted conversation snapshots to a bounded optional disk cache
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
- #181 Windows: the engine, vision encoder and MCP servers end with the server (closing the window no longer orphans them)
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)
- #175 Preserve independent conversation caches across interleaved requests
- #169 generate/verify: make diagnostic dumps work under --spec for native packs
- #167 generate: zero session state for native packs before the prompt path
- #158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUs
- #154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)
- #153 Expert cache keeps ~4 GiB more on WDDM (+10% decode on a 32 GB card); Ctrl+C and client hang-ups on Windows
- #151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setup
- #140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S support
- #136 Linux prebuilt from CI: CUDA 12.8, sm_80/86/89 + PTX, attached to each release
- #130 setup: Volta (V100, sm_70) and Turing as an experimental build path
- #129 core: add optional shared expert arena backing
- #126 Windows: the expert load runs 24x slower from Task Scheduler (0.05 vs 1.42 GiB/s); a hint names the cause
- #124 Pascal (sm_60): bring the engine up on compute capability 6.0
- #111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU
- #105 The Monitor's GPU readings on ROCm: an amdgpu sysfs backend, and the disk rate without psutil
- #102 feat: optional RAM and VRAM guard for the server it launches
- #100 fix: honor native per-layer layouts in file-backed expert loading
- #91 GGUF reader: refuse a duplicate tensor name at open, naming the tensor and the file
- #89 loader: bulk-read the expert arena and the MTP drafter (MSVC splits ifstream reads into 4095-byte calls)
- #88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)
- #87 Turing port: run on RTX 20 (sm_75) and newer
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)
- #79 linux: make DirectFile reads actually asynchronous
- #63 Fix Windows PermissionError in get_llama_cpp()
- #59 Fix/53 sampler penalties. sampler: penalties apply once, as in llama.cpp's default chain (#53); guard the unsized penalty bitmap
- #42 pinned: MEM_LARGE_PAGES on Windows never worked - enable the privilege, round the size
- #41 serve: preserve primary conversation state across auxiliary calls
- #37 core: elastic expert cache via CUDA VMM and on-demand GPU vision lifecycle
- #22 tools: add real-time web dashboard for hardware and engine monitoring
- #19 serve: per-request temperature/top_p/top_k/min_p/penalties sampling
- #16 Add optional multi-GPU expert caches and compact transfers
- #14 CPU pool: fix dangling else in Linux physical_cores() (fixes #13, root cause of #11)
- #10 serve: read short prompt parts through the decode windows (~1 s faster first token)
- #9 CPU pool: sleep between requests instead of spinning (fixes #4)
- #8 Conversation cache: keep the chat between requests, read only what is new