トピック / Benchmarks
Benchmarks
Auto topic: Benchmarks
Issues
- #1484 [SYCL] 2x Arc Pro B70 + UD-IQ4_XS: the port's default expert kernels give corrupted output (silent wrong answers); STRATA_EXPERT_SPLIT=1 restores it
- #1483 Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm building
- #1479 HIP: STRATA_HEAD_MIX_MULTI / STRATA_ONE_TOKEN_COMMIT are hard-off; as opt-ins head-mix is exact and +0.4% decode on R9700 (patch inside)
- #1478 gfx1201 (R9700): the gfx1151 prompt switches in arch_defaults.cpp are exact and +4-8% prompt; suggest a gfx12 entry (decode half is slower)
- #1473 STRATA_SYCL_SPIN_MAX default: a JIT build of the B70 gets the A-series bound and loses 46-67% of decode
- #1469 0.1.40.2: decode on a Tesla P40 (sm_61) is about half of 0.1.40's; bisected to 2e4ddf6, and restoring __restrict__ in pdl.hpp (STRATA_PDL_RESTRICT) gives back most of it
- #1463 MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)
- #1453 Live agent workload on RX 7900 XTX (24 GB) + 30 GB RAM — 70–74 tok/s, 91–96% expert cache hits
- #1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combine
- #1440 [Bug]: [SYCL] 0.1.40.2 on 2x Arc Pro B70: --layer-split fails three ways (#1054's fix from #1111 is not in main, plus two new split regressions)
- #1428 --layer-split auto on a P40 + RTX 3070 (3070 first) flips from K=43 to K=47 with --kv int8 or --vram-reserve-mib 600, and prefill drops from 458 to 111-113 tok/s (0.1.40.2)
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
- #1412 Windows: `SeLockMemoryPrivilege` moves the expert arena to 2 MiB large pages — but the token is issued at logon (1314 ≠ 1450)
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying
- #1409 What is normal prefill tok/s? Suggest to publish this info
- #1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request
- #1404 [Feature]: run the image encoder on a second PC
- #1397 SYCL port on Windows/OpenCL: 0.1.40-sycl decodes at 22.3 tok/s with --spec 2 where 0.1.39-sycl does 53.6 (and short runs vary 40%)
- #1389 RX 7900 XTX (gfx1100): 925 -> 2,503 tok/s prompt read (+171%) - the method, and the 26% cliff below full lend coverage
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by default
- #1375 Strata on an RX 6600 XT (gfx1032, 8 GB) and a Xeon E5-2678 v3 with DDR3
- #1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it
- #1369 Two 32K prompts on one RTX 5090: overlapping them matched back-to-back wall time
- #1353 stager-transient-regression
- #1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40
- #1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stack
- #1347 serve: a full conversation cache drops the snapshot at the physical-RAM gate instead of evicting to pass it (270k tokens re-read, 96 s, 32 times)
- #1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the model
- #1337 setup --calibrate loses the measurements it already has when a later engine start fails
- #1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)
- #1302 0.1.40 SYCL port on 2x Arc Pro B65: expert-cache fill stalls (1-core spin + xe "Timedout job" + device coredump); 0.1.38-based engine works on the same card
- #1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
- #1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
- #1286 PCIe Issue with Strata and RTX 3060
- #1285 Dual RTX 5090: Strata 0.1.40 production decode at 243 tok/s and synthetic peer-prefill results
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1277 HIP: fused IQ expert kernels via RDNA4 int8 WMMA (RX 9060 XT / gfx1200)
- #1272 0.1.40: with the shared-expert stream fork off by default, decode on 2x RX 6900 XT (gfx1030, Linux) is 5-13% slower; STRATA_SH_STREAM=1 restores it, and the #884 stall does not come back with it
- #1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speed
- #1261 461
- #1258 gfx1151 (Strix Halo) prompt: per-kernel profile of a 16K prompt (rocprofv3) - where would outside help be useful?
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
- #1252 DraftPolicy: a lookup window size priced during a slow stretch stays over-priced for the rest of the process
- #1239 v0.1.40 two-GPU tester results on a P40 + RTX 3070: 150-request soak clean, --pipeline-windows 2 +9% / +15% decode in 5 of 5 pairs
- #1238 0.1.40: --layer-split auto with STRATA_STAGE_TRIM=1 still picks K=2 on a P40 + RTX 3070; 572b607 fixes it but is not in the release
- #1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32
- #1234 Windows/RTX 4090 (0.1.39): stall report stage "decode -1" for 61 s; 0 layers served - thread stacks attached
- #1233 ``` Windows + Strix Halo (gfx1151, Radeon 8060S): UD-IQ4_XS runs end-to-end — decode flat 29.5-30.4 tok/s at 1K-32K, but a GPU driver timeout (TDR) during a cold 32K prefill ```
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1211 Bench: --ple-io direct vs. --ple-io ram on system with sufficient headroom of RAM
- #1200 RTX 5060 Ti 16 GB / Windows / IQ3_S: 36–42 tok/s — tuning directions & method (local gate scan, large pages, PLE FP8)
- #1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores it
- #1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)
- #1168 Client-supplied stop sequences are ignored (OpenAI `stop` and Anthropic `stop_sequences`) — blocks agent/harness use
- #1165 Ship an MMQ-enabled engine build (or opt-in loading) for k-quant prompts: a full-RAM machine is GPU-bound on the FP16 dequant path
- #1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)
- #1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
- #1139 GDN batch path skips activation quantization when STRATA_QFUSE=1
- #1138 Windows: higher process/thread priority improves decode throughput under CPU contention, with a measured background-work trade-off
- #1128 0.1.40 --kv-grow: a short request trims the K/V to 8192 cells and drops the conversation cache, so the next 128K+ continuation is re-read in full (20.9% of continuations, RTX 5090, IQ3_S)
- #1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML type
- #1118 Batch slots: the engine stalls ("no progress for 60 s … reading the prompt (batched)") when a long prompt is read while another slot decodes; STRATA_BATCH_DECODE_SHARE=0 avoids it
- #1113 sycl: the v0.1.40.1 port does not build as shipped (3 compile errors, 12 undefined symbols)
- #1103 hip: gfx1030 (RX 6800 XT) intermittent verify timeouts (#267): ROCm 7.14 gfx103X wheels regression, fixed by 7.13
- #1094 --layer-split auto fails on dual GPUs for 512k context
- #1085 Performance regression after updating to v0.1.40 - RTX 5090
- #1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
- #1059 serve: a request the engine rejects with ERR waits the full 300 s before failing (STOP + _drain_control("DONE") after an ERR), holding the control lock
- #1056 Prefill: Stager threads yield-spin while waiting. On Linux with --mmap-experts (DGX Spark / GB10), prompts are read ~10x slower; elsewhere ~20 cores are burned during prefill
- #1054 [SYCL / Intel Arc] 2x Arc Pro B70 --layer-split auto deadlocks at startup after the host-mirror fill
- #1053 A reply that ends inside `<think>` (no `</think>`) comes back empty with `finish_reason: "stop"`, and the agent stops — close the thinking and continue once
- #1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
- #1018 [gfx906] gr_up_fast_kernel<GrMulti> aborts the engine with HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION on long-context chats (intermittent; the non-fast path is stable)
- #1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"
- #1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)
- #997 0.1.39: engine exits (code 1) with "parallel": 4 on UD-IQ4_XS, RTX 3090 24 GB, three concurrent 8K prompts (reproducible; IQ3_S unaffected)
- #980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missing
- #979 HIP: a gfx1201 hipBLASLt table for 1.2.1 (100201, ROCm 7.2.0) makes the prompt 6-17% slower than no table
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
- #968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
- #954 [BUG] 0.1.39: fused int8 prompt experts (STRATA_PF_FUSED=1) + #583 byte-budget ring → kernel hang, `prefill mmq: iota: unknown error` on sm_120 (RTX 5090, WSL2)
- #953 Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decode
- #951 crash after first inputs
- #942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39
- #941 UPD: Successfully run on RTX 3090 24gb + Z590E + 64gb ddr4 3200 2ch
- #937 `[gfx1030] HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION in gdn_step_commit_kernel after a few long fresh prefills`
- #922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runs
- #921 CPU expert-pool 20 ms spin can hurt GPU-heavy decode; 100 µs gives ~15% higher TG and ~4× lower CPU usage on dual RTX 4090
- #915 Windows AMD report: RX 6800M (gfx1031, laptop) works with a self-built HIP engine — please add 0x73DF / gfx1031 to the Windows path
- #909 Proposal (docs only): a short "contributing a change or report" section for AGENTS.md, and a pinned list of test requests. Yes/no?
- #906 Older PC (Haswell, DDR3, PCIe 3.0 x8): +15% decode on Q2_0 and IQ3_S from #706, #764 and a busier adaptive tier
- #892 Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works
- #890 HIP: "parallel": 2 + --layer-split on 2x R9700 exits (code 1) after "captured the batch window over slots 0,1"
- #875 P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)
- #871 Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)
- #870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
- #867 Intel SYCL port on 2x Arc Pro B60: results, and "never rang (graph finished)" on the host-mirror path
- #843 Tool calls silently dropped in agent sessions after an empty assistant turn (fix: skip empty assistant turns when rendering)
- #841 not stable start with 3+ cards
- #840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaningless
- #825 Results for --model UD-IQ4_XS
- #816 0.1.39: decode on an RX 6800 (gfx1030, Windows) is 12-22% slower than 0.1.38; STRATA_SH_STREAM=0 removes it
- #798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`
- #791 Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split
- #781 1M context: identical expert work per window, but the VRAM-touching stages grow 7-8x
- #779 laptop crash with latest
- #776 --batch on a 100%-resident layer split: engine dies at "captured the batch window over slots 0,1" (zero-doorbell path?)
- #775 Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile
- #746 Regarding 128GB Memory Optimization Strategies
- #737 [enhancement] #498 follow-up: UD-Q4_K_XL split gate — allow confirm/override just below 135 GB total RAM
- #729 KV int8 vs fp16 KLD compare at 240K context
- #728 Infinite loop issue
- #723 Design question before we build: two lanes on a peer pair that help each other ("mutual help") — welcome upstream, or keep it in a fork?
- #713 Proposal: a common benchmark format and a comparison page for community reports
- #703 [AMD HIP] RX 6750 XT (gfx1031) works in heterogeneous layer split; long multimodal prompts can crash the engine
- #697 Windows HIP / gfx1201: prompt stalls at `GPU rang 0`; post-store `__threadfence_system()` resolves it
- #692 Feature Request: Intel Arc A770 Support
- #691 Windows: generation slows when the server window is minimized (EcoQoS moves threads to E-cores) — fix included
- #690 Layer split: native prefill kernels fail on the second GPU (no kernel image available), regardless of which card it is
- #684 [Question] Prompt reProcessing
- #675 Feature request: `logprobs` / `top_logprobs` on `/v1/chat/completions` (typed decisions from one forward pass)
- #669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores
- #658 Proposal: single-GPU prefill buffer planning and RAM budgeting for conversation caching (RTX 5090 measurements)
- #642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)
- #641 AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64): a working port, numbers, and PRs to upstream it
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
- #616 RTX A3000 / 0.1.38: AVX2 i-quant gate/up codebook gather — bit-exact ~1.3x isolated, +1.0% e2e, kept as opt-in
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #610 RTX A3000 / 0.1.38: a native IQ3_XXS decode window decomposes via STRATA_VERIFY_PROFILE, and its GPU side is memory-bandwidth-bound
- #606 After a 36,689-token request at 155k, every later request answers one repeated token until the engine reloads
- #605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)
- #604 Resident expert placement on a layer split, for mixed or older GPUs: where I think it helps, plus some doc notes, and an offer to test
- #602 Measured numbers: Qwen3.8-Flash-Next on 2× RTX 3060 (Q2_0 vs IQ3_S), plus a local reproduction of #75
- #597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090
- #590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 W
- #588 Reported expert-cache hit rate excludes --pcie-frac experts, so it rises as tok/s falls
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)
- #577 0.1.38: UD-Q4_K_XL prompts 15-40% slower on a 96 GB PC, from the unbuffered file tier (STRATA_UNBUFFERED_LOAD=0 restores them)
- #560 AMD/Linux desktop: default VRAM reserve (700 MiB) lets the driver evict ~24 GB to RAM, OOM kills kwin; --vram-reserve-mib 3072 fixes it
- #556 EXL3 (exllamav3 trellis) weight support
- #548 decode_cluster_parity: graph replays race their input upload (pageable cudaMemcpy + non-blocking stream)
- #542 sm_120 + IQ packs: ~4x prefill regression 0.1.34-0.1.36 when libcudart's ABI is mismatched - the #420 fits() gate silently disables MMQ (smpbo reads 1)
- #541 AMD/HIP (R9700, engine 0.1.36): batched prompt read stalls at a repeatable token; STRATA_PLE_BATCH=0 avoids it
- #535 Thinking speed suddenly drops to 0.5 token/s
- #528 Conversation cache: decode drops ~4-6x on continued conversations (0.1.36, Windows, RTX 5090, IQ3_XXS)
- #519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt path
- #511 Engine stopped unexpectedly
- #506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)
Pull requests
- #1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegister
- #1497 docs: tool selection and recovery analysis with model-family guide
- #1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)
- #1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMA
- #1493 Experimental cooperative memory relief and resumable text requests
- #1491 bench: completed Flash Next tool-use results with charts and examples
- #1486 feat: introduce Dockerfile.rocm-stable
- #1472 Intel Arc A770 (DG2, 16 GB) support in the SYCL port
- #1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3
- #1462 bench: RTX 5090 Laptop GPU follow-up for 0.1.40.2 and 0.1.40.3
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K
- #1459 bench: RX 7900 XTX + 6800 XT helper layout, 1M and long-output results
- #1458 helper cache: configurable VRAM allowance with allocation-floor checks
- #1456 bench: Flash Next variant tool-eval campaign
- #1452 bench: community report, Tesla V100 32 GB + P100 16 GB expert-helper (mixed Volta/Pascal), Flash-Next IQ3_XXS
- #1449 Add Metal backend support for Strata on macOS
- #1441 Prompt chunks sized for a layer split's pipeline (STRATA_PREFILL_PIPE=1, opt-in)
- #1439 perf(prefill): cache immutable BF16-to-FP16 weight conversions
- #1437 perf(cuda): add opt-in shape-tuned sm75 interleaved verify
- #1436 fix(serve): isolate CPU cores for independent Linux replicas
- #1435 docs: report RTX 8000 measurements of existing runtime switches
- #1433 community bench: RTX 5090 + RTX PRO 4000 Blackwell, UD-Q4_K_XL helper cache, engine 0.1.40.3
- #1432 bench: Arc Pro B65 Gen4 results on Strata v0.1.40.2
- #1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262K
- #1424 Pascal GP100: a Q8_0 decode GEMV (STRATA_Q8_SM60=1, opt-in) and IQ3_XXS / IQ3_S tables in shared memory
- #1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)
- #1419 bench: replay token IDs and verified windows in serve decode
- #1418 cuda: opt-in MMVQ activation reuse across output rows
- #1415 cpu: add bit-exact AVX2 singleton expert rows
- #1411 OLDER_GPUS.md: the P100 row now has a measurement (2x P100, IQ3_S)
- #1408 docs: a short "Contributing a change or report" section in AGENTS.md and a test-requests list
- #1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweep
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1401 V100 (sm_70): prompt experts on FP16 tensor cores (+10% prompt) and three existing decode kernels as defaults (-6.8% GPU per window)
- #1395 hip: a hipBLASLt tuning table for gfx1150 (Radeon 890M), and gfx1151's exact speed switches on gfx1150
- #1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
- #1382 docs: index entry for the TITAN RTX community report
- #1381 mtp: the draft layer's experts at Q8_0 and BF16, alongside Q2_0
- #1378 bench: community report, RX 6600 8 GB (gfx1032), EPYC 7232P, IQ3_XXS 128K on Linux
- #1377 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.2
- #1376 bench: add RX 7900 XT results for IQ3_S and IQ3_XXS
- #1374 experts: IQ1_S on GPU/AVX-2, IQ1_S+IQ1_M on AVX-512, and a measured PCIe share
- #1372 DeltaNet recurrence: STRATA_GDN_CHUNKED=1 - the prompt's recurrence in 32-token chunks (opt-in, 1.8x faster, other bits)
- #1370 prefill: opt-in STRATA_HC_UPMIX=1 on CUDA - the hyper-connection up projection with gr_mix_r as its epilogue (sm_80+)
- #1368 prefill: quantize each token's MoE input once and scatter it to its k rows (the same bytes)
- #1361 tools/hip: gfx1151 hipBLASLt 100500 tuning table (Windows ready-made …
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)
- #1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4
- #1354 docs: the expert cache is a budget, and on Windows an over-sized one pages instead of failing
- #1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #1345 calibrate: the sweep visits every value in every round instead of once
- #1339 cuda: add an opt-in SM120 Q8 T8 QKV MMQ path
- #1333 bench: dual RTX 5090 production and synthetic Strata 0.1.40 results
- #1332 calibrate: the PCIe share sweep reaches 1.0, where no expert goes to the CPU pool
- #1330 Add community benchmark: 2x Tesla P100 16GB
- #1325 file tier (Windows): read each expert from experts.bin and a mirror on a second drive at once
- #1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)
- #1323 prefill: batched unbuffered stager reads; close the experts.bin view once reads are unbuffered (port of #833)
- #1321 bench: publish an IQ2_XS speed report for the Radeon AI PRO R9700
- #1320 test: qualify existing exact HIP paths on gfx906
- #1319 serve: add opt-in bounded Responses history and continuation
- #1318 docs(AMD_HIP): my R9700 numbers for hipBLASLt 1.4.1 and STRATA_HIP_WMMA
- #1316 spec: re-probe stale full chains when shorter chains pay
- #1315 glm5-next: GLM-5.3-Flash support {intial attempt & prefill really sucks avoid testing its not ready }
- #1309 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
- #1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #1296 hip: chunk-parallel GDN recurrence (WY form), two parity-gated arms (STRATA_GDN_WY)
- #1295 Experimental native Windows path for Intel Arc
- #1293 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly
- #1292 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)
- #1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context
- #1290 Fix shared MTP slot initialization and speculative row bounds
- #1289 hip: the gfx1100 100401 table covers the small-T expert GEMM (prompt +31.9%)
- #1287 feat(vision): support Windows HIP image encoding
- #1282 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
- #1279 perf(cuda): use exact GP100 VMAD for native DP4A emulation
- #1278 HIP: run the K-quant MMQ test on HIP builds; UD-Q4_K_XL measured on an R9700
- #1270 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.40.1 - one card, layer split and expert helper; stock, the two gfx103x switches, and #1150 #1151 #1167
- #1263 Community benchmark: RTX 5090 Laptop GPU (24 GB), Windows 11, Strata 0.1.40.1
- #1262 STRATA_HC_REQ8: hyper-connection projections requantized to int8 + fp32 scale per 32 at load (rebased #704)
- #1259 setup: --draft-vocab it - the English/code subset plus the tokens of Italian text
- #1253 Enable serial multi-GPU batch MTP
- #1251 --peer-device prompt share without P2P: host route, per-token sums, FP16 transfers (2.1-2.5x prompts on a no-P2P pair)
- #1249 Pipelined batch: two requests run at half speed (pad rows in group windows, unbalanced slot choice)
- #1245 bench: UD-IQ4_XS first NVIDIA measurement (RTX 5090 Laptop, 24 GB)
- #1243 Bench/2026 10 05 community rtx 5080
- #1240 Avoid unused MTP branch stream on HIP
- #1237 perf(expert): the cache fill's stage buffers pinned with cudaHostAlloc
- #1226 Community benchmark: 2x RTX PRO 4500 Blackwell, Swift IQ3_XXS, engine 0.1.40.1 (solo to 256k, --batch after #776)
- #1225 bench: add R9 dual RTX 2080 Ti community results
- #1223 generate: allow the elastic K/V with --peer-device (VMM ranges keep their memory on the owning GPU)
- #1219 setup: a Spanish draft subset (--draft-vocab es); Spanish answers dra…
- #1216 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens (reland)
- #1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoring
- #1204 serve: a Chinese / English language switch for the web app
- #1199 Community benchmark: RTX 3090 eGPU + 64GB, IQ2_XS vs IQ3_XXS
- #1193 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #1192 Community benchmark: 4x Tesla P100 16 GB (Pascal), IQ3_XXS, CUDA 12 engine
- #1186 setup: --draft-vocab es - the English/code subset plus the tokens of Spanish text
- #1184 AMD HIP: the primary ordinal, and the helper caches beside the resident RAM mode
- #1181 perf(expert): the verify window's GPU plan in O(n) instead of O(n^2)
- #1179 batch: with --batch-groups, a group's lowest free slot first (+14-22% at 2-3 clients)
- #1173 bench: community report, RX 7900 XTX on Windows 11, IQ3_S, engine 0.1.40
- #1171 docs: minor prefill config change results in faster prefill
- #1167 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)
- #1166 cpu pool: --host-core sibling, the host off the interrupts' logical processor (hybrid CPUs too)
- #1159 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1158 bench: community reports, RTX 4090, IQ3_XXS 204800 and IQ3_S 143360 - 0.1.38 to 0.1.40.1
- #1157 Community benchmark: 2x Tesla P100, Flash-Next IQ3_S, 128k context
- #1154 decode: support a three-GPU two-window pipeline on 3 P4s
- #1151 hip: the FP32 tiled QSA block scorer on HIP cards (opt-in STRATA_SELECT_SIMT=1, as on CUDA; a 128K prompt +7.4% on gfx1030)
- #1150 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)
- #1149 HIP gfx103x FP16 prompt (#835 follow-up): STRATA_F16_RANGE measures what reaches FP16's range, and hip_prefill_gemm covers the FP16-io route
- #1142 hip: accept gfx1010 and gfx1011 (RDNA1) in the arch gate
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
- #1137 Community results: Strata 0.1.40 Q4/Q8, MTP and ngram on RTX PRO 6000
- #1134 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload (re-submission of #466)
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
- #1131 generate: the prompt loan keeps its residency bookkeeping without the token graph (#796 part B)
- #1125 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
- #1124 Community results: Strata 0.1.40 Q4/Q8 on RTX PRO 6000
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)
- #1122 Pipelined windows + async adaptive tier together (--pipeline-windows 2 --adapt-async 1)
- #1120 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, opt-in)
- #1115 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #1112 web: make MCP tool activity readable and collapsible
- #1111 Intel arc 0.1.40
- #1109 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
- #1107 prefill: gather experts in groups on short prompts too
- #1104 docs: visual community RTX PRO 6000 concurrency results (64K/128K)
- #1102 fix: batch host paths honor refreshed residency (deadlocks the engine during a prompt loan)
- #1101 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)
- #1100 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)
- #1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)
- #1097 V100: bit-exact GEMV / MoE operator fusions (experimental build only)
- #1096 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.2% e2e)
- #1095 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
- #1091 qsa_select: FP32 tiled block scores below sm_80 (+22% 128K / +44% 248K prefill on a 2080 Ti)
- #1087 cpu: allow configuring expert-pool tasks per phase
- #1084 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
- #1081 fix(build): qualify size_t in vmm.hpp (GCC 12 rejects the unqualified name)
- #1078 bench: RX 6800M (gfx1031) community report - self-built HIP engine on Windows
- #1077 feat(arm): build and run Strata on aarch64, including DGX Spark
- #1067 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
- #1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer split
- #1057 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)
- #1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)
- #1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #1038 decode: STRATA_PCIE_BALANCE=1 picks each layer's PCIe count from measured costs (opt-in)
- #1033 prefill: gather GPU-resident experts in groups on short prompts too
- #1030 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in, CUDA VMM)
- #1028 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer split
- #1027 bench: community report, RX 6700 XT 12 GB (gfx1031), Ryzen 7 5700G, I…
- #1025 fix: unblock all-resident batch windows (batch verify waited for a doorbell the graph never publishes)
- #1024 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)
- #1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)
- #1021 V100: bit-exact GEMV / MoE operator fusions (experimental build only)
- #1020 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.5% e2e)
- #1019 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)
- #1017 feature: RDNA3 support / Cache-aware routing
- #1016 Bench: RTX 5060 Ti 16 GB on Ryzen 9 7945HX, IQ3_S
- #1011 kv: one pinned KV pool shared by the sessions (--kv-pool-tokens; 4 lanes x 256K pin 6.24 GiB instead of 15.5)
- #1010 batch: up to 3 MTP drafts per slot inside the 8-row window (2 lanes, code: 71.6 -> 84.3 tok/s)
- #1007 hip: the QSA block scores on rocBLAS SGEMM, the solution measured on the card (opt-in; a 128K prompt +2.8% on gfx1030)
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)
- #1003 resident RAM mode: release each layer's mapped experts as it is copied (Windows)
- #1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)
- #1001 generate: wait for the residency table's upload before the next window (#871)
- #1000 HIP: gfx1102 (RX 7600) community-validated, with an end-to-end run (follow-up to #192)
- #996 hip: accept gfx1010 and gfx1011 (RDNA1) in the arch gate
- #995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)
- #993 bench: community report, RTX 5070 Ti + Ryzen 9 5900X, 64 GB DDR4-3600 (IQ2_XS / IQ3_XXS / IQ3_S, six-arm A/B)
- #988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)
- #981 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)
- #969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)
- #966 serve: recover tool calls the model writes next to the template's form, and return the rest as content
- #965 serve: pin Claude Code's per-request billing stamp, so the conversation cache reuses agent prompts
- #955 bench: Arc Pro B65 Gen4 results on patched v0.1.40
- #950 prefill: opt-in pointer-list bypass for gate/up gather (STRATA_PF_PTR_MMQ)
- #949 cpu: allow configuring expert-pool tasks per phase
- #944 hip: gfx11 WMMA prompt attention and selection scorer, gfx1151 hipBLASLt table
- #943 Add experimental deepMoE Vulkan backend for DeepSeek text chat
- #936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…
- #934 Per-layer expert cache: each layer's slots are that layer's blob
- #933 CPU trellis kernels (AVX2 / AVX-512) for TQ2_T / TQK6 / TQK7 expert-cache misses (stacked on #928)
- #930 cpu: IQ3_S / IQ2_S AVX2 decode without the per-index shift/mask/or (1.07-1.6x per row)
- #928 Gyro (TQ2_T / TQK6 / TQK7 + Hadamard rotation): native support for the agentionai Qwen3.8-Flash-Next quants
- #927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854
- #924 serve: render Codex's compaction with the conversation's tool prefix (Codex compatibility)
- #923 serve: answer Codex's thread-title requests without running them (Codex compatibility)
- #919 hip: Windows support for Radeon gfx1151 APUs (on #895)
- #917 bench: community report, Qwen3.8-Flash-Next-GSQ-RCO Q2_0 (Strix Halo/Radeon 8060S gfx1151 + 128 GB VRAM)
- #913 bench: community report, RTX 4060 Ti + Xeon E5-2673 v3 (DDR3, PCIe 3.0 x8), Q2_0 and IQ3_S, 0.1.39 path vs #706 + #764 + a busier tier
- #912 setup: show the GPU's PCIe link in step 1, warn on fewer lanes than the card has
- #910 Two-GPU decode: bounded attention merge (bitwise, opt-in) and pipelined PLE read-ahead / early chain (stacked on #905)
- #907 calibrate: measure the adaptive expert tier (--adapt-every / --adapt-swaps / --adapt-decay)
- #905 Pipelined windows + async adaptive tier together (stacked on #859 and #876)
- #904 Verify window: PDL on sm_90+, graph branches, batched Q4_0 KV append and PLE post-ops (bit-identical)
- #903 prefill: a prompt reads in equal chunks; a short streaming tail stays (#693, rebased)
- #902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)
- #901 Add ARM64 NVIDIA GB10 build and native-pack support
- #898 perf: Q8 resident adaptation (+24.5–28.2% rotation gain on RTX PRO 6000)
- #895 hip: support Radeon gfx1151 APUs
- #889 sycl: read doorbell flags with an uncached L1+L3 hint (the GPU never saw the host's store on an Arc Pro B60)
- #888 hip: chunk-parallel GDN recurrence (WY form), two parity-gated arms - opt-in STRATA_GDN_WY (measured on gfx1100)
- #887 layer split: score every four-way placement instead of guessing it
- #885 Perf/duplex transfers only
- #883 sm_60: PTX vmad for the __dp4a fallback (bit-exact, ~2.2x in isolation)
- #882 bench: community report, Flash-Next IQ2_XS vs a dense 27B as coding agents (RTX 4090 + 32 GB RAM)
- #880 layer split: the stage weight trim works under `auto` too 4X gain in the prefill in 4 way ring
- #876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)
- #869 serve: opt-in recovery from repeated reasoning passages
- #868 docs: Arc Pro B60 rows and notes, and the B60's PCI id (e211) in setup_intel.py
- #866 sycl: keep cudaStreamQuery's answer in the per-layer ring waits
- #865 ngram: add opt-in Q8_0 PLE table support
- #864 perf: opt-in resident expert exchange buffer rotation
- #863 cpu kernels: gathered decode on the cores where it is faster, AVX-VNNI rows, IQ3_S one-token kernel
- #860 Unsloth UD-Q6_K_XL: Q8_0 PLE rows and Q6_K gate/up experts
- #859 Layer split: pipelined verify windows for one conversation (--pipeline-windows, opt-in)
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
- #856 hip: RDNA3 (gfx1100) WMMA kernel for the QSA block-score select/scorer (opt-in STRATA_SELECT_WMMA, gfx12/gfx11)
- #854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)
- #853 cuda: specialize hyper-connection up for short verify windows
- #851 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
- #850 bench: community report, Tesla V100 32 GB + RTX 4070 layer split with 16 GB of RAM (IQ2_XS)
- #849 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)
- #848 Resident RAM mode with a layer split, --vram-reserve-later-mib (#642)
- #846 Keep MTP drafting in concurrent batch slots
- #844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)
- #837 server: "return_progress": true puts the prompt's PP line on the stream as llama.cpp's prompt_progress
- #835 HIP gfx103x: the prompt path's 16-bit GEMMs in FP16 in and out (prompt reads ~2x on RX 6900 XT)
- #834 bench: community report, RTX 4090 + 32 GB RAM at 512K context (IQ2_XS)
- #833 Prefill: batched unbuffered expert reads, close the experts.bin view; opt-in RAM budget with helper caches
- #832 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens
- #823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)
- #820 prefill: the dense GGUF projections through MMQ on HIP (opt-in STRATA_DENSE_MMQ=1)
- #818 Windows: opt-in release of mapped expert pages after GPU uploads
- #815 bench: community results for an RX 6800 on Windows (0.1.39 against 0.1.38)
- #811 Add an RTX 5090 + Ryzen 7 9700X IQ3_S benchmark on engine 0.1.39
- #809 Intel arc 0.1.39 + perf fixes
- #808 hip/gfx906: cudaFuncSetAttribute must be a function, not a macro (ROCm 7.2.1 build fix)
- #802 Run IQ1_S (Unsloth UD-IQ1_S) on the GPU and AVX-2, and read the expert file tier faster
- #800 qsa_prompt_attn: Volta (sm_70) m8n8k4 kernel - keep the hi and lo halves in separate accumulator chains
- #799 docs: the expert cache size is a budget, and on Windows an over-sized one pages instead of failing
- #797 sycl: enable Arc A770 inference and document measured performance
- #796 generate: one startup VRAM plan before the expert cache is committed (#765)
- #794 split: the auto placement sees each stage's PCIe link
- #792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layer
- #789 prefill: overlap shared-expert work with routed expert uploads
- #786 hip: RDNA3 (gfx1100) WMMA kernel for the int8-KV QSA prompt attention (opt-in STRATA_HIP_WMMA, gfx12/gfx11)
- #783 perf(cuda): fused decode/verify/MTP kernels and graph launch reductions on 0.1.39
- #780 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #777 Add an RTX 5090 + Ryzen 7 9700X IQ3_XXS benchmark on engine 0.1.39
- #773 file tier: unbuffered reads on Linux too (O_DIRECT + the kernel's asynchronous reads)
- #766 hip: a gfx1100 hipBLASLt tuning table for 100401 (packaged ROCm 10.0.0)
- #763 serve: enforce OpenAI tool_choice and allow structured output with tools
- #762 serve: extract JSON from prose and code fences
- #761 stage no longer waits for the whole split below it
- #758 bench: community MI50 (gfx906) — Strata 0.1.38 speed and recall
- #757 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
- #755 hip: gfx1100 hipBLASLt 1.5.0 tuning table (ROCm 10.2 nightly) — closes the gfx1100 100500 gap
- #751 Persist conversation KV state in a bounded disk LRU cache
- #745 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly, serve-path decode (Qwen3.8-Flash-Next IQ3_S)
- #744 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #743 qsa_select: coalesced streaming top-k for long Turing prompts (+23% 248K prefill)
- #742 qsa_select: FP32 tiled block scores below sm_80 (+22% 128K / +33% 248K prefill on a 2080 Ti)
- #741 prefill: BF16 projections on FP16 tensor cores below sm_80 (+8-17% prefill on a 2080 Ti)
- #740 Community benchmark: RX 7900 GRE (gfx1100, 16 GB), three Flash-Next quants
- #734 Reuse long chat history when the last user message is edited
- #733 Pipeline CPU expert work without a batch-wide phase barrier
- #732 Overlap streamed KV uploads with prefill computation
- #731 Avoid duplicate expert slots across helper and primary GPUs
- #726 Add opt-in live cache resizing and desktop resource presets
- #721 Community benchmark: 2× RTX 3060 12 GB, Qwen3.8-Flash-Next IQ3_S (three prompt sizes, 3 runs each)
- #717 fix: make the CUDA language link the shared runtime (Windows sm_70 LNK1169, #585)
- #707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)
- #704 STRATA_HC_Q8: hyper-connection projections as int8 + fp32 scale per 32
- #699 Linux: read ahead at startup (cold start ~920 s → 70 s)
- #698 Bench: RTX 5090 / 9950X3D (IQ3_S, three-pass A/B) + RTX 2080 Ti / 9900KF (IQ3_XXS, Turing)
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)
- #693 prefill: auto:16384 tries every 1024 tokens above 8192, and equal chunks - prompts 21-38% faster from 20K on a 16 GB card
- #688 Bench: RTX 5090 / Ryzen 9 9950X3D, IQ3_S, v0.1.34 vs v0.1.38 vs fused prompt kernels
- #682 serve: a Chinese / English language switch for the web app
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #677 bench: Linux + 2x AMD Instinct MI50 16 GB (gfx906), Coder IQ1_M, 4K to 128K prompt tokens
- #674 bench: community report, RTX 3090 Ti + 2x Xeon E5-2699 v3 (HP Z840), IQ3_S at 262K
- #667 Add Intel XPU backend: Qwen decode on Arc Pro B60 via SYCL
- #664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)
- #663 prefill: a layer split's idle card streams a share of a one-chunk prompt's experts
- #655 sm_75: run the prompt path's BF16 products on the FP16 tensor cores (+15-18 % prefill on RTX 2080 Ti)
- #651 ple: read Q8_0 and Q5_1 n-gram tables
- #650 pinned: back the Linux expert arena with transparent huge pages
- #646 perf(cuda,decode): zero-doorbell resident verify graph, sub-warp expert packing & shared-mem staging (+39-72% tok/s)
- #639 layer split: with explicit split points each GPU loads only its own layers' dense weights
- #638 hip: AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64) as an opt-in build
- #628 bench: report Coder IQ1_M on Threadripper 3990X and RX 6900 XT
- #627 opt for sm_70(just test V100-16g*1 or 2)
- #626 cpu: preserve full thread affinity on Windows and Linux
- #624 bench: community report — RTX 4090, Ryzen 9 7950X, Flash-Next IQ3_S
- #622 iq: gather the AVX2 codebook from the grid table (STRATA_IQ256_GATHER)
- #618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokens
- #603 qsa_select: a 1,024-thread top-k past the register kernel's reach (CUDA) - a 243K-token prompt +26% on an RTX 3060
- #600 qsa_prompt_attn: tensor-core kernel for Volta (sm_70, mma.m8n8k4)
- #599 native experts: Q4_0 and Q4_1 GPU kernels, Q4_0 PLE table
- #593 prefill: run BF16 GEMMs through FP16 tensor cores on sm_70/sm_75
- #586 ple: read the PLE table at full BF16 precision (opt-in, alongside FP8)
- #583 Prefill chunk ring
- #578 Improve multi-GPU decode performance via optimized expert-helper execution
- #568 parity tests: ple_parity runs without the missing fixtures, gr_parity and ple_parity wait for their setup
- #565 tools/hip: add gfx1100 hipBLASLt tuning table for hipBLASLt 1.2.2 (100202)
- #563 Give the expert cache's VRAM back without unloading the model (#533)
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
- #540 volta (sm_70) and turing (sm_75): faster prompt reading - BF16 projections on the FP16 tensor cores, and a faster pre-Turing attention kernel
- #538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokens
- #527 Experimental P100: pinned RAM complement for fixed layer-split CPU/GPU inference
- #521 bench: add community RTX 4000 Ada 2-GPU vs 3-GPU layer-split results
- #512 prefill: use active QSA top-k bounds on Turing with large KV capacity
- #508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35
- #501 verify: pin the serve host thread the session loop already pins
- #500 perf(cpu-pool): eliminate host serialization barrier for intermediate activation quantization with adaptive threshold
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
- #489 bench/results: RTX A3000 12GB laptop + i7-12850HX (sm_86): hugepages, PCIe probe, calibration
- #488 pinned: honor STRATA_NO_LARGEPAGES on Linux; name the hugetlb pool shortfall
- #483 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34
- #482 Networked Pool support
- #473 native experts: Q5_0 GPU kernel (#222)
- #469 bench: community results, RTX 2080 Ti 11 GB + Threadripper 3960X, IQ3_S at 262K with KV streaming
- #466 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload
- #464 ple: read the n-gram table at full BF16 precision
- #463 decode: wait for an adaptive expert swap before reading the residency table
- #462 diagnostics: dump every verify window's logits, so greedy decode can be bisected
- #453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-resident
- #452 qsa_prompt_attn: Q4_0 KV on tensor cores (mode 4) - long prompts with --kv q4_0 ~19% faster
- #443 kernels: add opt-in small-window GR and one-warp MMVQ experiments
- #442 hip: add community gfx1012 support with version-gated legacy compatibility
- #441 Contrib/non mtp serving
- #440 bench: community report, RTX 5090 + Ryzen 9 9950X3D, IQ3_S at the full 262K window; prefill auto:32768 A/B; conversation-cache 32 GiB demo
- #439 prefill: a group's expert gathers in one launch (bit-identical, long prompts +4%)
- #436 serve: a non-streaming request stops when its client disconnects (#430, part of #431)
- #433 bench: community report, RTX 5090 + Ryzen 9 5950X (AVX2), UD-Q4_K_XL / IQ3_S / Swift IQ3_XXS on engines 0.1.31-0.1.39
- #427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)
- #422 hip: enable experimental Strix Halo gfx1151 source builds
- #418 bench: community entry, 2x RTX PRO 4500, Swift IQ3_XXS, engine 0.1.30 and 0.1.36
- #417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31
- #416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budget
- #413 prefill: DeltaNet recurrence with the three value heads of a key head in one thread (bitwise; 1.4x on a 4080 SUPER, 1.3x on a 3090)
- #409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory
- #405 docs: add performance documentation and update repository hygiene
- #404 bench(gfx1100): 0.1.31 vs 0.1.30 prefill/decode, curated entry + raw trials
- #395 Feature/nvidia p40
- #394 AVX1: run on a CPU with AVX but no AVX2, with an end-to-end run on Sandy Bridge-E
- #389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #386 tools/hip: gfx1201 hipBLASLt table for ROCm 7.2.4 (hipBLASLt 1.2.2), ~1.8x prompt speed on the R9700
- #382 HIP: MTP prompt pass per group by default (fixes prompt hang on gfx1201 with KV streaming)
- #380 HIP: count the desktop's VRAM on Windows (WDDM budget, STRATA_WDDM_BUDGET)
- #378 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #376 HIP: a failed first configure no longer leaves the kernels at -O0
- #368 QSA select: RDNA3 (gfx1100) WMMA block-scores arm on #337's arch dispatch (opt-in STRATA_SELECT_WMMA=1)
- #366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Next
- #363 experts: smaller and fewer launches for the verify window's PCIe call
- #362 File tier: read the GGUF in place unbuffered when the file cache cannot keep it beside the RAM budget (Windows; #286 ported into the existing tier)
- #360 qsa_prompt_attn: Volta (sm_70) crashes in the prompt path - compare the full compute capability
- #353 NVFP4 routed experts on 0.1.35 (experimental): 120a only for the W4A4 unit (rebase of #292)
- #350 Chat: a context gauge, per-request timings, and conversation compaction
- #343 Don't pursue: BF16 decode GEMV via Turing tensor cores (sm_75) - measured negative
- #339 hip: hipBLASLt tuning table for gfx1201 (R9700) with ROCm 10.2.0a nightly (hipBLASLt 1.5.0)
- #337 hip: gfx12 (RDNA4) QSA select - WMMA block scorer, prompt-aware top-k dispatch
- #336 cuda: add experimental Tesla P4 8GB (sm_61) support
- #333 serve: add lazy startup and reliable Windows model cleanup
- #332 serve: add a standalone API request monitor
- #325 AMD working on Windows (9070XT tested)
- #323 hip: add experimental RX 5500 XT 8 GB (gfx1012) support
- #322 hip: fast packed-byte intrinsics for RDNA3/RDNA4 (v_perm_b32 + SWAR)
- #319 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)
- #318 iq_pack: store an F32 ple_conv1d as F16
- #315 hc: the multi-token hyper-connection read split finer
- #313 RDNA3 WMMA kernels for the gfx1100 prefill (opt-in, runtime-gated on gfx11; rebased on 0.1.30)
- #312 QSA select: ROCWMMA block-scores arm (opt-in STRATA_SELECT_WMMA=1, gfx1100)
- #311 HIP: support RDNA2 gfx1030 working and gfx103X is untested.
- #302 HIP/Windows: batch expert copy dependencies in long prefill
- #295 Pascal build runs (STRATA_EXPERIMENTAL_SM60)
- #294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windows
- #292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rows
- #291 PLE: the n-gram table in FP8 E4M3, as Qwen ships it (opt-in)
- #287 setup: --draft-vocab cyrillic (English/code + the whole Cyrillic script)
- #286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.bin
- #284 Verify: the commit graph no longer waits; it overlaps the MTP draft
- #283 Prompt path: BF16 projections get the activation's BF16 remainder too
- #282 prefill auto: chunks up to 32768 (was 8192), +15% on a 32K prompt
- #279 Windows: leave 1500 MiB of VRAM unused by default (decode stalled with 700)
- #278 serve: count_tokens, Anthropic thinking only when asked, STRATA_REQUEST_LINES, a relative exe
- #275 Add bounded disk persistence for evicted conversations
- #272 feat(cpu): add hybrid architecture awareness and --pool-affinity for P/E-core CPUs
- #270 qsa_prompt_attn: run the f16 tensor-core path on Turing (sm_75)
- #269 Split prefill: run the whole prompt on the main GPU (--prefill-main)
- #262 HIP: gfx1201 (Radeon AI PRO R9700) support, faster packed byte intrinsics, and a gfx1201 hipBLASLt table
- #260 RDNA3 WMMA kernels for the gfx1100 prefill (dense GEMMs and prompt attention)
- #257 q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact)
- #256 HIP: support RDNA4 gfx1200 (RX 9060 XT)
- #255 Q8 0 support
- #247 HIP: Windows and gfx1201 (RX 9070)
- #242 IQ kernels: decode each weight part once for every column and entry, bitwise identical
- #241 Faster grouped and per-hit Q2_0 expert kernels, bitwise identical
- #237 Use a second GPU's VRAM as an opt-in expert store for the prompt path
- #234 Bench: RTX 3090 / dual RTX 3090 on EPYC 7453
- #233 docs: add community benchmark guide and RTX 5090 results
- #229 multi-GPU: second GPU as an opt-in expert-cache tier with P2P rows (--peer-device), part 1 of #204
- #228 QSA select on tensor cores (rocWMMA, opt-in, gfx1100) + AMD RX 7900 XTX rates row + verify profiler columns fix
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)
- #216 Multi-GPU: the session carve and the per-stage prompt loans
- #207 E-2 on the AVX-2 path too: prefetch the i-quant expert rows (same switch, CPUs without AVX-512)
- #192 HIP: support RDNA3 gfx1101/gfx1102 (RX 7600/7700/7800), not only gfx1100
- #190 Spill evicted conversation snapshots to a bounded optional disk cache
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
- #185 sampler: the sampled path over 64 vocabulary chunks per token (identical picks, ~12% faster sampled decode)
- #177 Web UI: prefill speeds, request timings, a context gauge, chat compaction
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)
- #175 Preserve independent conversation caches across interleaved requests
- #154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)
- #153 Expert cache keeps ~4 GiB more on WDDM (+10% decode on a 32 GB card); Ctrl+C and client hang-ups on Windows
- #148 tests: bitwise check of native_mmvq's multi_exact contract, with a negative control
- #139 perf(prefill): add shape-selected pre-Ampere BF16 SGEMM with current-main A/B
- #136 Linux prebuilt from CI: CUDA 12.8, sm_80/86/89 + PTX, attached to each release
- #130 setup: Volta (V100, sm_70) and Turing as an experimental build path
- #124 Pascal (sm_60): bring the engine up on compute capability 6.0
- #121 Add gfx1100 HIP support with tuned prefill and bounded expert uploads
- #111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU
- #110 Multi-GPU: a second GPU as an expert tier for decode and prompts (--peer-device), + multi-conversation cache
- #109 decode: batch the verify window's per-token kernels (bit-identical, +~10% decode)
- #108 prefill: bit-identical kernel speed-ups (batched indexer append, parallel GDN conv, column-split GDN recurrence, one-launch embedding gather)
- #104 The Monitor's live tok/s is a rate over the last seconds, not the mean since the first token
- #101 build: allow sm_70 (Volta) builds — the only sm_80+ dependency is unreferenced
- #94 Add experimental gfx1100 HIP backend for RX 7900 XTX
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)
- #79 linux: make DirectFile reads actually asynchronous
- #67 Support OrcaRouter IQ3_XXS with explicit BF16 compatibility packing
- #62 serve: the conversation cache keeps its shared prefix; the rest rotates LRU
- #54 Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)
- #43 kernels: AVX2 multi-token i-quant row kernels (an AVX2-only CPU decoded every i-quant row once per token)
- #22 tools: add real-time web dashboard for hardware and engine monitoring
- #19 serve: per-request temperature/top_p/top_k/min_p/penalties sampling
- #14 CPU pool: fix dangling else in Linux physical_cores() (fixes #13, root cause of #11)
- #12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…
- #10 serve: read short prompt parts through the decode windows (~1 s faster first token)
- #9 CPU pool: sleep between requests instead of spinning (fixes #4)
- #8 Conversation cache: keep the chat between requests, read only what is new
- #7 serve: fix zero-token replies from engine-queue race between requests