Temas / Documentation
Documentation
Auto topic: Documentation
Issues
- #1488 prefill.cpp: a GCC 10 build fails (std::atomic::wait is a GCC 11 library feature, with no guard and no documented minimum)
- #1484 [SYCL] 2x Arc Pro B70 + UD-IQ4_XS: the port's default expert kernels give corrupted output (silent wrong answers); STRATA_EXPERT_SPLIT=1 restores it
- #1483 Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm building
- #1482 [Feature]: publish a ready-made Docker image, so users do not compile the engine themselves
- #1473 STRATA_SYCL_SPIN_MAX default: a JIT build of the B70 gets the A-series bound and loses 46-67% of decode
- #1464 Add vision to docker container
- #1463 MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)
- #1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combine
- #1444 can kvcache be saved and reused on disk?
- #1440 [Bug]: [SYCL] 0.1.40.2 on 2x Arc Pro B70: --layer-split fails three ways (#1054's fix from #1111 is not in main, plus two new split regressions)
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying
- #1404 [Feature]: run the image encoder on a second PC
- #1397 SYCL port on Windows/OpenCL: 0.1.40-sycl decodes at 22.3 tok/s with --spec 2 where 0.1.39-sycl does 53.6 (and short runs vary 40%)
- #1389 RX 7900 XTX (gfx1100): 925 -> 2,503 tok/s prompt read (+171%) - the method, and the 26% cliff below full lend coverage
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by default
- #1375 Strata on an RX 6600 XT (gfx1032, 8 GB) and a Xeon E5-2678 v3 with DDR3
- #1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it
- #1353 stager-transient-regression
- #1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)
- #1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the model
- #1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
- #1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
- #1299 AMD vison encoder on Windows
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speed
- #1258 gfx1151 (Strix Halo) prompt: per-kernel profile of a 16K prompt (rocprofv3) - where would outside help be useful?
- #1248 SYCL build fails on Intel Arc A770 (Alchemist) - v0.1.40.1
- #1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32
- #1234 Windows/RTX 4090 (0.1.39): stall report stage "decode -1" for 61 s; 0 layers served - thread stacks attached
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1211 Bench: --ple-io direct vs. --ple-io ram on system with sufficient headroom of RAM
- #1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???
- #1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores it
- #1191 Automatic support of other variants and quants
- #1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)
- #1165 Ship an MMQ-enabled engine build (or opt-in loading) for k-quant prompts: a full-RAM machine is GPU-bound on the FP16 dequant path
- #1161 [Feature Request] Port ngram-mod & other ngram-* self-speculative techniques?
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
- #1135 k8v4: deterministic token repetition on sequential generation (RTX 5090, 0.1.40) — byte-identical across two quants
- #1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML type
- #1113 sycl: the v0.1.40.1 port does not build as shipped (3 compile errors, 12 undefined symbols)
- #1074 Build fails on openEuler 24.03 / GCC 12.3 (sm_70): vmm.hpp uses an unqualified size_t
- #1071 vmm.cpp: the CUDA 12 engine does not build with a toolkit older than 12.5 (missing CUDART_VERSION guard)
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
- #1054 [SYCL / Intel Arc] 2x Arc Pro B70 --layer-split auto deadlocks at startup after the host-mirror fill
- #1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)
- #1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
- #1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)
- #1005 CRLF line-break tokens (`\r\n`, ids 317/845/8301) get far too little probability; everything else matches llama.cpp
- #990 MCP install planning rejects supported AMD CPU vision (and retains the old Windows HIP restriction)
- #977 --check prints [ok] for below-floor RAM and a misleading 'This PC can run Strata' verdict
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
- #968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly
- #942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39
- #932 Dual GPU support
- #925 Setup: show progress in the browser from the first second (status page with steps, speed and ETA)
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
- #909 Proposal (docs only): a short "contributing a change or report" section for AGENTS.md, and a pinned list of test requests. Yes/no?
- #884 HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)
- #881 [Windows / AMD] RX 6800M (gfx1031) runs with a self-built HIP engine — plus 2 Windows bugs found (packaging encoding, vision `vcvars`)
- #870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
- #841 not stable start with 3+ cards
- #840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaningless
- #831 The profile's length caps the expert arena — even an explicit --expert-cache N
- #803 Perplexity ~7-10% above llama.cpp on the same GGUF (Qwen3.8-Flash-Next), also at 2K context, every version tested
- #795 Main engine strata.exe crashes with STATUS_ILLEGAL_INSTRUCTION (0xC000001D) on non-AVX-512 CPU - 20 WER crashes across v0.1.35/v0.1.38, fault offsets cluster in fixed code regions
- #775 Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile
- #772 Community project: strata-router, a multi-node router for Strata / 社区项目:Strata 多节点路由入口
- #767 --vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL grounding
- #760 524K via rope-scaling on consumer Blackwell (sm_120, no clusters): the histogram fallback works end to end — field data + two budget gotchas on 16GB cards
- #756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)
- #747 Feature request: multi-conversation parking on multi-GPU layer-split
- #729 KV int8 vs fp16 KLD compare at 240K context
- #713 Proposal: a common benchmark format and a comparison page for community reports
- #675 Feature request: `logprobs` / `top_logprobs` on `/v1/chat/completions` (typed decisions from one forward pass)
- #658 Proposal: single-GPU prefill buffer planning and RAM budgeting for conversation caching (RTX 5090 measurements)
- #644 Specifying explicit GPU split crashes Strata
- #642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)
- #641 AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64): a working port, numbers, and PRs to upstream it
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)
- #604 Resident expert placement on a layer split, for mixed or older GPUs: where I think it helps, plus some doc notes, and an offer to test
- #602 Measured numbers: Qwen3.8-Flash-Next on 2× RTX 3060 (Q2_0 vs IQ3_S), plus a local reproduction of #75
- #601 Linux build fails on Ubuntu 26.04 (glibc 2.43) with CUDA 12.9 (nvcc host-math conflict)
- #590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 W
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)
- #564 Gui settings
- #551 Wrong UI on Windows?
- #544 Please run several security audits
- #543 Can we add OpenCode config.jsonc to docs?
- #519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt path
- #516 desktop (KDE/Wayland): compositor VRAM exhaustion when expert-cache auto fills a 24 GB card - pin framebuffer -12, KWin graphics reset
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)
Pull requests
- #1499 core: write reduced helper rows directly to mapped output
- #1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegister
- #1497 docs: tool selection and recovery analysis with model-family guide
- #1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)
- #1493 Experimental cooperative memory relief and resumable text requests
- #1490 Add a workflow that builds and pushes the Docker image on a version tag
- #1489 Responses persistence and disk-only cache: CUDA/HIP integration
- #1486 feat: introduce Dockerfile.rocm-stable
- #1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)
- #1477 docs: correct pipeline async compatibility (#1447)
- #1472 Intel Arc A770 (DG2, 16 GB) support in the SYCL port
- #1471 feat: optionally offload idle session state while keeping model weights loaded
- #1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3
- #1462 bench: RTX 5090 Laptop GPU follow-up for 0.1.40.2 and 0.1.40.3
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K
- #1459 bench: RX 7900 XTX + 6800 XT helper layout, 1M and long-output results
- #1452 bench: community report, Tesla V100 32 GB + P100 16 GB expert-helper (mixed Volta/Pascal), Flash-Next IQ3_XXS
- #1449 Add Metal backend support for Strata on macOS
- #1448 #606 follow-up: clamp the three remaining unclamped q8_1 scale/sum emit sites
- #1446 Responses: replay fixes, experimental persistence and optional summaries
- #1441 Prompt chunks sized for a layer split's pipeline (STRATA_PREFILL_PIPE=1, opt-in)
- #1439 perf(prefill): cache immutable BF16-to-FP16 weight conversions
- #1438 perf(peer): add opt-in byte-capacity-aware expert placement
- #1437 perf(cuda): add opt-in shape-tuned sm75 interleaved verify
- #1436 fix(serve): isolate CPU cores for independent Linux replicas
- #1435 docs: report RTX 8000 measurements of existing runtime switches
- #1434 telemetry: GPU load, VRAM, temperature and power for AMD cards on Windows (#1380)
- #1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262K
- #1426 dflash: experimental standalone DeepSpec DFlash block drafter for Qwen3.8-Flash-Next (greedy, correctness-first)
- #1424 Pascal GP100: a Q8_0 decode GEMV (STRATA_Q8_SM60=1, opt-in) and IQ3_XXS / IQ3_S tables in shared memory
- #1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)
- #1420 Add optional GLM-5.3-Flash support from Project Maya
- #1415 cpu: add bit-exact AVX2 singleton expert rows
- #1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer split
- #1408 docs: a short "Contributing a change or report" section in AGENTS.md and a test-requests list
- #1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweep
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1401 V100 (sm_70): prompt experts on FP16 tensor cores (+10% prompt) and three existing decode kernels as defaults (-6.8% GPU per window)
- #1400 docs: running Strata behind Open WebUI
- #1398 docs: keeping Codex CLI off the network with Strata
- #1396 hip/gfx906: the tree does not build (q6_k MMQ instance and cudaEventBlockingSync)
- #1395 hip: a hipBLASLt tuning table for gfx1150 (Radeon 890M), and gfx1151's exact speed switches on gfx1150
- #1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
- #1388 hip: gfx1151 hipBLASLt tuning table for setup's ROCm 7.14.0a20260608 (hipBLASLt 1.4.0)
- #1382 docs: index entry for the TITAN RTX community report
- #1378 bench: community report, RX 6600 8 GB (gfx1032), EPYC 7232P, IQ3_XXS 128K on Linux
- #1374 experts: IQ1_S on GPU/AVX-2, IQ1_S+IQ1_M on AVX-512, and a measured PCIe share
- #1361 tools/hip: gfx1151 hipBLASLt 100500 tuning table (Windows ready-made …
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)
- #1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4
- #1354 docs: the expert cache is a budget, and on Windows an over-sized one pages instead of failing
- #1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #1343 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers
- #1340 core: retain fixed RAM originals during native async exchanges
- #1339 cuda: add an opt-in SM120 Q8 T8 QKV MMQ path
- #1338 cuda: reuse secondary Q8 copies in native async refills
- #1336 cuda: compact traversal for read-only Q8 miss fills
- #1335 docs: supporting QFUSE one-token commit evidence for #1209
- #1330 Add community benchmark: 2x Tesla P100 16GB
- #1326 serve: support IPv6 bind addresses
- #1321 bench: publish an IQ2_XS speed report for the Radeon AI PRO R9700
- #1318 docs(AMD_HIP): my R9700 numbers for hipBLASLt 1.4.1 and STRATA_HIP_WMMA
- #1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #1295 Experimental native Windows path for Intel Arc
- #1293 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly
- #1292 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)
- #1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context
- #1288 serve: a long read gives way by what is left to read, not prompt length (#656)
- #1287 feat(vision): support Windows HIP image encoding
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
- #1279 perf(cuda): use exact GP100 VMAD for native DP4A emulation
- #1278 HIP: run the K-quant MMQ test on HIP builds; UD-Q4_K_XL measured on an R9700
- #1270 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.40.1 - one card, layer split and expert helper; stock, the two gfx103x switches, and #1150 #1151 #1167
- #1266 Add the Update the engine card
- #1265 Add the engine updater: verify, stage, all-or-nothing swap, rollback
- #1262 STRATA_HC_REQ8: hyper-connection projections requantized to int8 + fp32 scale per 32 at load (rebased #704)
- #1260 tests: order MMQ parity transfers on the compute stream
- #1245 bench: UD-IQ4_XS first NVIDIA measurement (RTX 5090 Laptop, 24 GB)
- #1241 serve: the Responses API reads Codex's additional_tools input items as tools (#782)
- #1227 mcp: the install plan follows setup's AMD rules - Windows AMD and vision=cpu on Linux are planned (#990)
- #1226 Community benchmark: 2x RTX PRO 4500 Blackwell, Swift IQ3_XXS, engine 0.1.40.1 (solo to 256k, --batch after #776)
- #1222 serve: render a Codex compaction with that conversation's tool prefix
- #1221 serve: answer Codex thread-title turns without reading them
- #1218 Check the ready-made engine against GitHub's SHA-256 before installing it
- #1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoring
- #1202 dashboard: Serve the web page under an optional dashboard path
- #1199 Community benchmark: RTX 3090 eGPU + 64GB, IQ2_XS vs IQ3_XXS
- #1197 docs: --calibrate loads the model more than once, and starts it when …
- #1193 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #1192 Community benchmark: 4x Tesla P100 16 GB (Pascal), IQ3_XXS, CUDA 12 engine
- #1186 setup: --draft-vocab es - the English/code subset plus the tokens of Spanish text
- #1184 AMD HIP: the primary ordinal, and the helper caches beside the resident RAM mode
- #1173 bench: community report, RX 7900 XTX on Windows 11, IQ3_S, engine 0.1.40
- #1172 serve: opt-in recovery of tool calls the model writes next to the template's form ("tool_call_recovery": true)
- #1171 docs: minor prefill config change results in faster prefill
- #1163 conversation cache: a re-parked conversation reserves at most 1/8 of each K/V buffer instead of up to 16 MiB, so fewer get evicted
- #1159 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1157 Community benchmark: 2x Tesla P100, Flash-Next IQ3_S, 128k context
- #1154 decode: support a three-GPU two-window pipeline on 3 P4s
- #1152 Fix Windows AMD telemetry and add optional Arabic UI
- #1149 HIP gfx103x FP16 prompt (#835 follow-up): STRATA_F16_RANGE measures what reaches FP16's range, and hip_prefill_gemm covers the FP16-io route
- #1147 docs: translate README to Italian
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
- #1134 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload (re-submission of #466)
- #1132 docs: running Strata behind llama-swap
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)
- #1119 Add local PDF, Word and Excel attachments with screenshot previews
- #1117 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in; resubmission of #1030 on the new main)
- #1115 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #1114 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1111 Intel arc 0.1.40
- #1109 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
- #1104 docs: visual community RTX PRO 6000 concurrency results (64K/128K)
- #1095 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
- #1092 launcher: pick a model and its settings in one window (experimental)
- #1090 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #1086 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
- #1078 bench: RX 6800M (gfx1031) community report - self-built HIP engine on Windows
- #1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer split
- #1060 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1045 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)
- #1038 decode: STRATA_PCIE_BALANCE=1 picks each layer's PCIe count from measured costs (opt-in)
- #1036 docs: --calibrate loads the model more than once, and starts it when …
- #1032 serve: add opt-in bounded Responses history and continuation
- #1030 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in, CUDA VMM)
- #1028 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1019 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
- #1017 feature: RDNA3 support / Cache-aware routing
- #1016 Bench: RTX 5060 Ti 16 GB on Ryzen 9 7945HX, IQ3_S
- #1014 vision: downscale large images before CPU encoding (12MP: >8 min → 20 s)
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)
- #1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)
- #1000 HIP: gfx1102 (RX 7600) community-validated, with an end-to-end run (follow-up to #192)
- #995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)
- #993 bench: community report, RTX 5070 Ti + Ryzen 9 5900X, 64 GB DDR4-3600 (IQ2_XS / IQ3_XXS / IQ3_S, six-arm A/B)
- #988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)
- #987 serve: /props reports chat_template_caps so Zed offers tools (#986)
- #981 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)
- #969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)
- #966 serve: recover tool calls the model writes next to the template's form, and return the rest as content
- #965 serve: pin Claude Code's per-request billing stamp, so the conversation cache reuses agent prompts
- #960 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #947 fix: avoid CPU doorbell waits in fully resident batch decoding
- #946 launcher: pick a model and its settings in one window (experimental)
- #943 Add experimental deepMoE Vulkan backend for DeepSeek text chat
- #936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…
- #934 Per-layer expert cache: each layer's slots are that layer's blob
- #933 CPU trellis kernels (AVX2 / AVX-512) for TQ2_T / TQK6 / TQK7 expert-cache misses (stacked on #928)
- #928 Gyro (TQ2_T / TQK6 / TQK7 + Hadamard rotation): native support for the agentionai Qwen3.8-Flash-Next quants
- #927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854
- #924 serve: render Codex's compaction with the conversation's tool prefix (Codex compatibility)
- #923 serve: answer Codex's thread-title requests without running them (Codex compatibility)
- #920 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
- #919 hip: Windows support for Radeon gfx1151 APUs (on #895)
- #913 bench: community report, RTX 4060 Ti + Xeon E5-2673 v3 (DDR3, PCIe 3.0 x8), Q2_0 and IQ3_S, 0.1.39 path vs #706 + #764 + a busier tier
- #902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)
- #898 perf: Q8 resident adaptation (+24.5–28.2% rotation gain on RTX PRO 6000)
- #885 Perf/duplex transfers only
- #882 bench: community report, Flash-Next IQ2_XS vs a dense 27B as coding agents (RTX 4090 + 32 GB RAM)
- #878 # Support sm_60, sm_70, and sm_86 with CUDA 12.6 or 12.8
- #868 docs: Arc Pro B60 rows and notes, and the B60's PCI id (e211) in setup_intel.py
- #865 ngram: add opt-in Q8_0 PLE table support
- #864 perf: opt-in resident expert exchange buffer rotation
- #861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)
- #860 Unsloth UD-Q6_K_XL: Q8_0 PLE rows and Q6_K gate/up experts
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
- #854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)
- #849 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)
- #844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)
- #839 Add experimental Linux gfx900 setup with wave64 validation
- #835 HIP gfx103x: the prompt path's 16-bit GEMMs in FP16 in and out (prompt reads ~2x on RX 6900 XT)
- #834 bench: community report, RTX 4090 + 32 GB RAM at 512K context (IQ2_XS)
- #832 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens
- #829 docs: running Strata behind llama-swap
- #827 docs: translate README to Italian
- #826 verify: the shared-expert stream fork starts off on HIP (fixes the 0.1.38 -> 0.1.39 decode regression, #816)
- #818 Windows: opt-in release of mapped expert pages after GPU uploads
- #815 bench: community results for an RX 6800 on Windows (0.1.39 against 0.1.38)
- #813 rtx6000fix
- #810 serve: accept json_schema roots that are unions of object schemas
- #809 Intel arc 0.1.39 + perf fixes
- #807 cuda: opt-in batched expert uploads to reduce host submission overhead
- #805 Lazy vision encoder: per-GPU elastic expert caches
- #800 qsa_prompt_attn: Volta (sm_70) m8n8k4 kernel - keep the hi and lo halves in separate accumulator chains
- #799 docs: the expert cache size is a budget, and on Windows an over-sized one pages instead of failing
- #797 sycl: enable Arc A770 inference and document measured performance
- #793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kit
- #784 sycl: follow #626 (ThreadAffinity) and #559 (NativeDense::load layer range) so v0.1.39 compiles
- #780 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #766 hip: a gfx1100 hipBLASLt tuning table for 100401 (packaged ROCm 10.0.0)
- #758 bench: community MI50 (gfx906) — Strata 0.1.38 speed and recall
- #755 hip: gfx1100 hipBLASLt 1.5.0 tuning table (ROCm 10.2 nightly) — closes the gfx1100 100500 gap
- #750 docs: describe a targeted ROCr USERPTR reclaim workaround
- #745 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly, serve-path decode (Qwen3.8-Flash-Next IQ3_S)
- #726 Add opt-in live cache resizing and desktop resource presets
- #721 Community benchmark: 2× RTX 3060 12 GB, Qwen3.8-Flash-Next IQ3_S (three prompt sizes, 3 runs each)
- #719 Add sandboxed HTML previews and scrollable chat code blocks
- #712 Fix Windows AMD telemetry and add optional Arabic UI
- #707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)
- #698 Bench: RTX 5090 / 9950X3D (IQ3_S, three-pass A/B) + RTX 2080 Ti / 9900KF (IQ3_XXS, Turing)
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)
- #695 hip: the tuning table's arch gate accepts the RDNA4 sibling (gfx1200 <-> gfx1201)
- #685 Add optional Windows HIP vision encoder builds and package validation
- #683 setup: a re-run of setup carries over the hand-edited config blocks (#629)
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #677 bench: Linux + 2x AMD Instinct MI50 16 GB (gfx906), Coder IQ1_M, 4K to 128K prompt tokens
- #674 bench: community report, RTX 3090 Ti + 2x Xeon E5-2699 v3 (HP Z840), IQ3_S at 262K
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #667 Add Intel XPU backend: Qwen decode on Arc Pro B60 via SYCL
- #666 setup, serve: the image encoder's device as its own role (--vision-device)
- #665 serve: no automatic --layer-split when the args use --peer-device
- #663 prefill: a layer split's idle card streams a share of a one-chunk prompt's experts
- #662 setup: prefer uv for the venv and package installs when uv is installed
- #660 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
- #653 serve: park conversations with a layer split
- #638 hip: AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64) as an opt-in build
- #636 native: add opt-in GBNF constraints to existing generation
- #630 serve: add opt-in stateless Responses API
- #619 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers
- #618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokens
- #614 serve: optionally checkpoint an existing chunk near the prompt tail
- #600 qsa_prompt_attn: tensor-core kernel for Volta (sm_70, mma.m8n8k4)
- #599 native experts: Q4_0 and Q4_1 GPU kernels, Q4_0 PLE table
- #571 Feat/responses api 451
- #565 tools/hip: add gfx1100 hipBLASLt tuning table for hipBLASLt 1.2.2 (100202)
- #563 Give the expert cache's VRAM back without unloading the model (#533)
- #562 serve: /slots reports n_prompt_tokens, so a llama.cpp-style context m…
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
- #554 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)
- #538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokens
- #531 multi-GPU: second GPU as an opt-in expert-cache tier (--peer-device), resubmit of #229 on v0.1.36
- #527 Experimental P100: pinned RAM complement for fixed layer-split CPU/GPU inference
- #512 prefill: use active QSA top-k bounds on Turing with large KV capacity
- #508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
- #492 Heterogeneous multi-GPU roles: whole model on one GPU, the MTP draft head on the other
- #491 docker: GGUF_DIR, RESIDENT_BUDGET_GIB and KV_STREAMING env vars; document the stop timeout
- #489 bench/results: RTX A3000 12GB laptop + i7-12850HX (sm_86): hugepages, PCIe probe, calibration
- #484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)
- #483 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #482 Networked Pool support
- #469 bench: community results, RTX 2080 Ti 11 GB + Threadripper 3960X, IQ3_S at 262K with KV streaming
- #466 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload
- #455 setup: mark the Unsloth family NVIDIA only so far in the AMD menu, and point a mixed PC to its NVIDIA card (#429)
- #454 serve: honor stop and stop_sequences
- #450 calibrate: a gpu list is not a layer split (#447)
- #443 kernels: add opt-in small-window GR and one-warp MMVQ experiments
- #442 hip: add community gfx1012 support with version-gated legacy compatibility
- #441 Contrib/non mtp serving
- #440 bench: community report, RTX 5090 + Ryzen 9 9950X3D, IQ3_S at the full 262K window; prefill auto:32768 A/B; conversation-cache 32 GiB demo
- #433 bench: community report, RTX 5090 + Ryzen 9 5950X (AVX2), UD-Q4_K_XL / IQ3_S / Swift IQ3_XXS on engines 0.1.31-0.1.39
- #427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)
- #426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)
- #422 hip: enable experimental Strix Halo gfx1151 source builds
- #418 bench: community entry, 2x RTX PRO 4500, Swift IQ3_XXS, engine 0.1.30 and 0.1.36
- #417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31
- #416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budget
- #409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory
- #407 Adaptive tier: --adapt-decay, and --adapt-tuned (opt-in): 31% fewer misses, -19% CPU pool, -22% PCIe on UD-Q4_K_XL
- #405 docs: add performance documentation and update repository hygiene
- #404 bench(gfx1100): 0.1.31 vs 0.1.30 prefill/decode, curated entry + raw trials
- #395 Feature/nvidia p40
- #389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #387 AMD: ship the gfx1200 hipBLASLt table, with the with/without measurements
- #386 tools/hip: gfx1201 hipBLASLt table for ROCm 7.2.4 (hipBLASLt 1.2.2), ~1.8x prompt speed on the R9700
- #366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Next
- #356 HIP: build the AMD engine on Windows
- #343 Don't pursue: BF16 decode GEMV via Turing tensor cores (sm_75) - measured negative
- #339 hip: hipBLASLt tuning table for gfx1201 (R9700) with ROCm 10.2.0a nightly (hipBLASLt 1.5.0)
- #336 cuda: add experimental Tesla P4 8GB (sm_61) support
- #334 serve: enforce structured JSON in the native sampler
- #330 V100 (sm_70): the compute-capability floor drops from 7.5 to 7.0
- #325 AMD working on Windows (9070XT tested)
- #323 hip: add experimental RX 5500 XT 8 GB (gfx1012) support
- #322 hip: fast packed-byte intrinsics for RDNA3/RDNA4 (v_perm_b32 + SWAR)
- #319 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)
- #313 RDNA3 WMMA kernels for the gfx1100 prefill (opt-in, runtime-gated on gfx11; rebased on 0.1.30)
- #311 HIP: support RDNA2 gfx1030 working and gfx103X is untested.
- #294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windows
- #292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rows
- #275 Add bounded disk persistence for evicted conversations
- #256 HIP: support RDNA4 gfx1200 (RX 9060 XT)
- #255 Q8 0 support
- #245 Doubled throughput
- #233 docs: add community benchmark guide and RTX 5090 results
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)
- #192 HIP: support RDNA3 gfx1101/gfx1102 (RX 7600/7700/7800), not only gfx1100
- #190 Spill evicted conversation snapshots to a bounded optional disk cache
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)
- #170 hip: gfx1100 hipBLASLt tuning table for ROCm 7.2.1 (100202)
- #169 generate/verify: make diagnostic dumps work under --spec for native packs
- #140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S support
- #130 setup: Volta (V100, sm_70) and Turing as an experimental build path
- #126 Windows: the expert load runs 24x slower from Task Scheduler (0.05 vs 1.42 GiB/s); a hint names the cause
- #121 Add gfx1100 HIP support with tuned prefill and bounded expert uploads
- #102 feat: optional RAM and VRAM guard for the server it launches
- #94 Add experimental gfx1100 HIP backend for RX 7900 XTX
- #87 Turing port: run on RTX 20 (sm_75) and newer
- #84 EXPERIMENTAL Rope scaling: contexts past the trained 262K (none / linear / YaRN)
- #81 Accept sm_80+ GPUs at runtime, matching the build guard and the README
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)
- #62 serve: the conversation cache keeps its shared prefix; the rest rotates LRU
- #59 Fix/53 sampler penalties. sampler: penalties apply once, as in llama.cpp's default chain (#53); guard the unsized penalty bitmap
- #52 nvme kv cache
- #41 serve: preserve primary conversation state across auxiliary calls
- #19 serve: per-request temperature/top_p/top_k/min_p/penalties sampling
- #12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…