标签 / Multi-GPU
Multi-GPU
Auto topic: Multi-GPU
Issues
- #1495 [Feature]: experimental per-stage KV-grow for a two-GPU layer split at native 262K
- #1488 prefill.cpp: a GCC 10 build fails (std::atomic::wait is a GCC 11 library feature, with no guard and no documented minimum)
- #1484 [SYCL] 2x Arc Pro B70 + UD-IQ4_XS: the port's default expert kernels give corrupted output (silent wrong answers); STRATA_EXPERT_SPLIT=1 restores it
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
- #1469 0.1.40.2: decode on a Tesla P40 (sm_61) is about half of 0.1.40's; bisected to 2e4ddf6, and restoring __restrict__ in pdl.hpp (STRATA_PDL_RESTRICT) gives back most of it
- #1468 prefill copy_i32: illegal memory access on long prompts when the prefill chunk is large (auto:32768)
- #1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combine
- #1440 [Bug]: [SYCL] 0.1.40.2 on 2x Arc Pro B70: --layer-split fails three ways (#1054's fix from #1111 is not in main, plus two new split regressions)
- #1428 --layer-split auto on a P40 + RTX 3070 (3070 first) flips from K=43 to K=47 with --kv int8 or --vram-reserve-mib 600, and prefill drops from 458 to 111-113 tok/s (0.1.40.2)
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying
- #1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request
- #1403 Almost no changes from 2 recent updates,
- #1387 [Feature Request] Support for higher quantizations (Q8 / unquantized) for Qwen 3.8 / 4 with high-end hardware (128GB VRAM / 192GB RAM)
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by default
- #1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40
- #1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stack
- #1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)
- #1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the model
- #1317 serve: an untimed read of the engine/vision pipe can hold the request turn forever (intermittent hang; the #481 lost-step class)
- #1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)
- #1304 Prefill/Reading the prompt with multi-gpu cards is slow
- #1302 0.1.40 SYCL port on 2x Arc Pro B65: expert-cache fill stalls (1-core spin + xe "Timedout job" + device coredump); 0.1.38-based engine works on the same card
- #1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
- #1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
- #1285 Dual RTX 5090: Strata 0.1.40 production decode at 243 tok/s and synthetic peer-prefill results
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1272 0.1.40: with the shared-expert stream fork off by default, decode on 2x RX 6900 XT (gfx1030, Linux) is 5-13% slower; STRATA_SH_STREAM=1 restores it, and the #884 stall does not come back with it
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
- #1239 v0.1.40 two-GPU tester results on a P40 + RTX 3070: 150-request soak clean, --pipeline-windows 2 +9% / +15% decode in 5 of 5 pairs
- #1238 0.1.40: --layer-split auto with STRATA_STAGE_TRIM=1 still picks K=2 on a P40 + RTX 3070; 572b607 fixes it but is not in the release
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1212 Review: shared MTP head fix — sibling Q4 state omission and test configuration gaps
- #1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)
- #1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1138 Windows: higher process/thread priority improves decode throughput under CPU contention, with a measured background-work trade-off
- #1094 --layer-split auto fails on dual GPUs for 512k context
- #1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40
- #1073 Wrench Thailand and Drill Bits Thailand: Reliable Hand and Power Tool Solutions from M10 TOOLS
- #1071 vmm.cpp: the CUDA 12 engine does not build with a toolkit older than 12.5 (missing CUDART_VERSION guard)
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
- #1054 [SYCL / Intel Arc] 2x Arc Pro B70 --layer-split auto deadlocks at startup after the host-mirror fill
- #1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)
- #1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
- #1018 [gfx906] gr_up_fast_kernel<GrMulti> aborts the engine with HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION on long-context chats (intermittent; the non-fast path is stable)
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
- #961 Windows (WDDM): verify window wedges nvlddmkm — 0x141 engine timeout → 0x116 TDR failure (driver defect, not a Strata bug)
- #952 BrokePipeError when terminating Strata on Linux
- #932 Dual GPU support
- #929 `--batch` with `--layer-split` exits ("verify batch: layer 34 never rang") with a 3070 first and a P40 second; works with the P40 first (CUDA, v0.1.39)
- #921 CPU expert-pool 20 ms spin can hurt GPU-heavy decode; 100 µs gives ~15% higher TG and ~4× lower CPU usage on dual RTX 4090
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
- #909 Proposal (docs only): a short "contributing a change or report" section for AGENTS.md, and a pinned list of test requests. Yes/no?
- #891 SYCL port v0.1.39 fails to compile: session.cpp ThreadAffinity type errors
- #890 HIP: "parallel": 2 + --layer-split on 2x R9700 exits (code 1) after "captured the batch window over slots 0,1"
- #884 HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)
- #875 P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)
- #870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
- #867 Intel SYCL port on 2x Arc Pro B60: results, and "never rang (graph finished)" on the host-mirror path
- #841 not stable start with 3+ cards
- #828 Heap corruption (munmap_chunk: invalid pointer) at startup in Prefill::init — crashes with 1/2/4 GPUs
- #776 --batch on a 100%-resident layer split: engine dies at "captured the batch window over slots 0,1" (zero-doorbell path?)
- #747 Feature request: multi-conversation parking on multi-GPU layer-split
- #737 [enhancement] #498 follow-up: UD-Q4_K_XL split gate — allow confirm/override just below 135 GB total RAM
- #713 Proposal: a common benchmark format and a comparison page for community reports
- #692 Feature Request: Intel Arc A770 Support
- #672 Better multiGPU support
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
- #647 Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS works on double Vega 20 GPU (Radeon Pro VII 16gb each).
- #644 Specifying explicit GPU split crashes Strata
- #641 AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64): a working port, numbers, and PRs to upstream it
- #629 Re-running setup with a different --context silently discards mcp_servers, mcp and sampling from the run config
- #623 avx1 not supprted for xeonV2
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
- #616 RTX A3000 / 0.1.38: AVX2 i-quant gate/up codebook gather — bit-exact ~1.3x isolated, +1.0% e2e, kept as opt-in
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)
- #557 AMD MultiGPU
- #530 high/xhigh reasoning_effort can silently run to max_tokens with empty content when reasoning_budget_tokens isn't set
- #520 [2x RTX PRO 6000] Would this approach allow me to run full bf16 model?
- #514 Feature Request: use AMD iGPU instead of CPU for offloading
- #509 layer_split auto is 2.86x slower than a balanced K on prefill with a heterogeneous pair (3080 + V100, measured)
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)
Pull requests
- #1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMA
- #1493 Experimental cooperative memory relief and resumable text requests
- #1490 Add a workflow that builds and pushes the Docker image on a version tag
- #1489 Responses persistence and disk-only cache: CUDA/HIP integration
- #1471 feat: optionally offload idle session state while keeping model weights loaded
- #1465 cuda: retain HC norm products in existing scratch
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1441 Prompt chunks sized for a layer split's pipeline (STRATA_PREFILL_PIPE=1, opt-in)
- #1436 fix(serve): isolate CPU cores for independent Linux replicas
- #1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262K
- #1426 dflash: experimental standalone DeepSpec DFlash block drafter for Qwen3.8-Flash-Next (greedy, correctness-first)
- #1424 Pascal GP100: a Q8_0 decode GEMV (STRATA_Q8_SM60=1, opt-in) and IQ3_XXS / IQ3_S tables in shared memory
- #1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)
- #1418 cuda: opt-in MMVQ activation reuse across output rows
- #1417 prefill: scope each split-stage loan estimate to its device
- #1411 OLDER_GPUS.md: the P100 row now has a measurement (2x P100, IQ3_S)
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1400 docs: running Strata behind Open WebUI
- #1382 docs: index entry for the TITAN RTX community report
- #1379 prefill: STRATA_PREFILL_CPU_SHARE=auto times layers with and without the share and shares only while that is faster (#1282)
- #1372 DeltaNet recurrence: STRATA_GDN_CHUNKED=1 - the prompt's recurrence in 32-token chunks (opt-in, 1.8x faster, other bits)
- #1370 prefill: opt-in STRATA_HC_UPMIX=1 on CUDA - the hyper-connection up projection with gr_mix_r as its epilogue (sm_80+)
- #1359 setup: link the engine against the toolkit whose nvcc builds it (CUDAToolkit_ROOT)
- #1358 build: vmm.cpp builds with a CUDA toolkit older than 12.5 again (#1071)
- #1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4
- #1354 docs: the expert cache is a budget, and on Windows an over-sized one pages instead of failing
- #1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #1333 bench: dual RTX 5090 production and synthetic Strata 0.1.40 results
- #1330 Add community benchmark: 2x Tesla P100 16GB
- #1325 file tier (Windows): read each expert from experts.bin and a mirror on a second drive at once
- #1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)
- #1323 prefill: batched unbuffered stager reads; close the experts.bin view once reads are unbuffered (port of #833)
- #1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context
- #1290 Fix shared MTP slot initialization and speculative row bounds
- #1284 Adaptive tier: --adapt-decay outside (0, 1) is refused (inf counts, NaN swap gains)
- #1271 serve: --conversation-cache-spill-dir keeps evicted conversations on disk across restarts
- #1270 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.40.1 - one card, layer split and expert helper; stock, the two gfx103x switches, and #1150 #1151 #1167
- #1259 setup: --draft-vocab it - the English/code subset plus the tokens of Italian text
- #1253 Enable serial multi-GPU batch MTP
- #1247 prefill: --prefill-help and --prefill-help-frac as the official flags for the idle-stage helper
- #1245 bench: UD-IQ4_XS first NVIDIA measurement (RTX 5090 Laptop, 24 GB)
- #1226 Community benchmark: 2x RTX PRO 4500 Blackwell, Swift IQ3_XXS, engine 0.1.40.1 (solo to 256k, --batch after #776)
- #1225 bench: add R9 dual RTX 2080 Ti community results
- #1223 generate: allow the elastic K/V with --peer-device (VMM ranges keep their memory on the owning GPU)
- #1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoring
- #1193 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #1192 Community benchmark: 4x Tesla P100 16 GB (Pascal), IQ3_XXS, CUDA 12 engine
- #1181 perf(expert): the verify window's GPU plan in O(n) instead of O(n^2)
- #1179 batch: with --batch-groups, a group's lowest free slot first (+14-22% at 2-3 clients)
- #1167 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)
- #1160 generate: --layer-split auto rejects a split that cannot start
- #1159 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1157 Community benchmark: 2x Tesla P100, Flash-Next IQ3_S, 128k context
- #1151 hip: the FP32 tiled QSA block scorer on HIP cards (opt-in STRATA_SELECT_SIMT=1, as on CUDA; a 128K prompt +7.4% on gfx1030)
- #1150 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)
- #1149 HIP gfx103x FP16 prompt (#835 follow-up): STRATA_F16_RANGE measures what reaches FP16's range, and hip_prefill_gemm covers the FP16-io route
- #1148 tests: hip_prefill_wmma_gemm_parity skips on a card without matrix cores instead of failing
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
- #1131 generate: the prompt loan keeps its residency bookkeeping without the token graph (#796 part B)
- #1120 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, opt-in)
- #1115 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #1114 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1111 Intel arc 0.1.40
- #1105 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vision footprint
- #1102 fix: batch host paths honor refreshed residency (deadlocks the engine during a prompt loan)
- #1100 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)
- #1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)
- #1098 prefill/gdn: the Volta prompt recurrence chain-split (+7-15%) and the opt-in chunked path
- #1097 V100: bit-exact GEMV / MoE operator fusions (experimental build only)
- #1096 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.2% e2e)
- #1095 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
- #1089 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing traces
- #1084 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
- #1083 hip/gfx906: the tree does not build (three STRATA_USE_HIP gates miss STRATA_HIP_GFX906)
- #1067 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
- #1066 hip/gfx906: the 0.1.40 tree does not build (three STRATA_USE_HIP gates miss STRATA_HIP_GFX906)
- #1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer split
- #1060 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)
- #1047 Adaptive tier: --adapt-decay outside (0, 1) is refused (inf counts, NaN swap gains)
- #1039 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #1028 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1025 fix: unblock all-resident batch windows (batch verify waited for a doorbell the graph never publishes)
- #1024 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)
- #1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)
- #1021 V100: bit-exact GEMV / MoE operator fusions (experimental build only)
- #1019 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
- #1017 feature: RDNA3 support / Cache-aware routing
- #1007 hip: the QSA block scores on rocBLAS SGEMM, the solution measured on the card (opt-in; a 128K prompt +2.8% on gfx1030)
- #995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)
- #991 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing traces
- #988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)
- #981 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)
- #949 cpu: allow configuring expert-pool tasks per phase
- #947 fix: avoid CPU doorbell waits in fully resident batch decoding
- #936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…
- #928 Gyro (TQ2_T / TQK6 / TQK7 + Hadamard rotation): native support for the agentionai Qwen3.8-Flash-Next quants
- #927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854
- #903 prefill: a prompt reads in equal chunks; a short streaming tail stays (#693, rebased)
- #897 setup: the disk check counts only what step 6 still writes
- #889 sycl: read doorbell flags with an uncached L1+L3 hint (the GPU never saw the host's store on an Arc Pro B60)
- #887 layer split: score every four-way placement instead of guessing it
- #885 Perf/duplex transfers only
- #883 sm_60: PTX vmad for the __dp4a fallback (bit-exact, ~2.2x in isolation)
- #882 bench: community report, Flash-Next IQ2_XS vs a dense 27B as coding agents (RTX 4090 + 32 GB RAM)
- #880 layer split: the stage weight trim works under `auto` too 4X gain in the prefill in 4 way ring
- #868 docs: Arc Pro B60 rows and notes, and the B60's PCI id (e211) in setup_intel.py
- #866 sycl: keep cudaStreamQuery's answer in the per-layer ring waits
- #854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)
- #852 engine: --gpu selects the visible GPUs without environment variables
- #850 bench: community report, Tesla V100 32 GB + RTX 4070 layer split with 16 GB of RAM (IQ2_XS)
- #849 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)
- #845 verify: batch windows honor the all-resident stage (no host doorbells)
- #835 HIP gfx103x: the prompt path's 16-bit GEMMs in FP16 in and out (prompt reads ~2x on RX 6900 XT)
- #823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)
- #818 Windows: opt-in release of mapped expert pages after GPU uploads
- #809 Intel arc 0.1.39 + perf fixes
- #807 cuda: opt-in batched expert uploads to reduce host submission overhead
- #805 Lazy vision encoder: per-GPU elastic expert caches
- #800 qsa_prompt_attn: Volta (sm_70) m8n8k4 kernel - keep the hi and lo halves in separate accumulator chains
- #796 generate: one startup VRAM plan before the expert cache is committed (#765)
- #794 split: the auto placement sees each stage's PCIe link
- #792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layer
- #789 prefill: overlap shared-expert work with routed expert uploads
- #783 perf(cuda): fused decode/verify/MTP kernels and graph launch reductions on 0.1.39
- #761 stage no longer waits for the whole split below it
- #734 Reuse long chat history when the last user message is edited
- #726 Add opt-in live cache resizing and desktop resource presets
- #718 serve: park conversations on disk, keyed by a client id; requests without one stay in RAM
- #716 conversation parking with --layer-split
- #711 kv: --kv k8v4 streams with --kv-resident
- #707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)
- #693 prefill: auto:16384 tries every 1024 tokens above 8192, and equal chunks - prompts 21-38% faster from 20K on a 16 GB card
- #677 bench: Linux + 2x AMD Instinct MI50 16 GB (gfx906), Coder IQ1_M, 4K to 128K prompt tokens
- #674 bench: community report, RTX 3090 Ti + 2x Xeon E5-2699 v3 (HP Z840), IQ3_S at 262K
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #666 setup, serve: the image encoder's device as its own role (--vision-device)
- #665 serve: no automatic --layer-split when the args use --peer-device
- #663 prefill: a layer split's idle card streams a share of a one-chunk prompt's experts
- #653 serve: park conversations with a layer split
- #652 serve: commit only the tokens a verify window hands out
- #651 ple: read Q8_0 and Q5_1 n-gram tables
- #650 pinned: back the Linux expert arena with transparent huge pages
- #646 perf(cuda,decode): zero-doorbell resident verify graph, sub-warp expert packing & shared-mem staging (+39-72% tok/s)
- #640 STRATA_ARENA_MMAP=1: the expert arena as a read-only mapped file (Linux), for small-RAM machines
- #639 layer split: with explicit split points each GPU loads only its own layers' dense weights
- #638 hip: AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64) as an opt-in build
- #637 serve: retry an engine start that fails, and keep the last known context after a failed restart
- #636 native: add opt-in GBNF constraints to existing generation
- #627 opt for sm_70(just test V100-16g*1 or 2)
- #622 iq: gather the AVX2 codebook from the grid table (STRATA_IQ256_GATHER)
- #614 serve: optionally checkpoint an existing chunk near the prompt tail
- #598 refill: improve multi-GPU processing speed ~27% PP improvement on 4× RTX 3060
- #591 Adaptive tier: --adapt-decay outside (0, 1) is refused (inf counts, NaN swap gains)
- #584 Split placement search
- #583 Prefill chunk ring
- #580 Weight carve 0138
- #578 Improve multi-GPU decode performance via optimized expert-helper execution
- #576 multi-GPU: conversation parking across a layer split
- #562 serve: /slots reports n_prompt_tokens, so a llama.cpp-style context m…
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
- #531 multi-GPU: second GPU as an opt-in expert-cache tier (--peer-device), resubmit of #229 on v0.1.36
- #527 Experimental P100: pinned RAM complement for fixed layer-split CPU/GPU inference
- #526 Allow conversation parking under --layer-split (owner-routed per-stage capture/restore)
- #521 bench: add community RTX 4000 Ada 2-GPU vs 3-GPU layer-split results
- #508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35
- #501 verify: pin the serve host thread the session loop already pins
- #492 Heterogeneous multi-GPU roles: whole model on one GPU, the MTP draft head on the other
- #483 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #482 Networked Pool support
- #450 calibrate: a gpu list is not a layer split (#447)
- #442 hip: add community gfx1012 support with version-gated legacy compatibility
- #418 bench: community entry, 2x RTX PRO 4500, Swift IQ3_XXS, engine 0.1.30 and 0.1.36
- #417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31
- #413 prefill: DeltaNet recurrence with the three value heads of a key head in one thread (bitwise; 1.4x on a 4080 SUPER, 1.3x on a 3090)
- #390 Multi-GPU: a --layer-split stage holds only its own layers' weights, and every card's leftover VRAM spills into a helper expert tier (4-way IQ3_S: 9,269 -> 14,172 expert slots, prefill chunk 2,048 -> 4,096)
- #389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Next
- #358 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #337 hip: gfx12 (RDNA4) QSA select - WMMA block scorer, prompt-aware top-k dispatch
- #320 verify: fix STRATA_VERIFY_PROFILE stage columns (unsigned stamp wrap + slot collision)
- #302 HIP/Windows: batch expert copy dependencies in long prefill
- #295 Pascal build runs (STRATA_EXPERIMENTAL_SM60)
- #294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windows
- #275 Add bounded disk persistence for evicted conversations
- #269 Split prefill: run the whole prompt on the main GPU (--prefill-main)
- #237 Use a second GPU's VRAM as an opt-in expert store for the prompt path
- #235 verify: STRATA_LOGPOS hook for per-position log-probabilities (items 12-13 of #84)
- #234 Bench: RTX 3090 / dual RTX 3090 on EPYC 7453
- #229 multi-GPU: second GPU as an opt-in expert-cache tier with P2P rows (--peer-device), part 1 of #204
- #227 AVX1: the Q2_0 expert kernels for CPUs that have AVX but no AVX2
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)
- #216 Multi-GPU: the session carve and the per-stage prompt loans
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)
- #203 serve: prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread)
- #202 ple: --ple-io ram locks the n-gram table in RAM (no SSD read on the prompt/token path)
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
- #188 prefill: software-pipelined loads in the column-split GDN recurrence (bit-identical)
- #187 decode: QSA block scores read each key block once per window (bit-identical)
- #186 decode: hyper-connection read in 2 kernels, stream-split (-1.5 ms per verify window)
- #185 sampler: the sampled path over 64 vocabulary chunks per token (identical picks, ~12% faster sampled decode)
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)
- #175 Preserve independent conversation caches across interleaved requests
- #169 generate/verify: make diagnostic dumps work under --spec for native packs
- #158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUs
- #151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setup
- #139 perf(prefill): add shape-selected pre-Ampere BF16 SGEMM with current-main A/B
- #129 core: add optional shared expert arena backing
- #121 Add gfx1100 HIP support with tuned prefill and bounded expert uploads
- #120 `--kv k8v4`: hybrid KV cache — INT8 K + Hadamard-rotated Q4_0 V (816 B/cell)
- #111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU
- #110 Multi-GPU: a second GPU as an expert tier for decode and prompts (--peer-device), + multi-conversation cache
- #109 decode: batch the verify window's per-token kernels (bit-identical, +~10% decode)
- #108 prefill: bit-identical kernel speed-ups (batched indexer append, parallel GDN conv, column-split GDN recurrence, one-launch embedding gather)
- #88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)
- #84 EXPERIMENTAL Rope scaling: contexts past the trained 262K (none / linear / YaRN)
- #22 tools: add real-time web dashboard for hardware and engine monitoring
- #16 Add optional multi-GPU expert caches and compact transfers
- #15 generate: residual control-vector projection from GGUFs