Themen / Linux
Linux
Auto topic: Linux
Issues
- #1479 HIP: STRATA_HEAD_MIX_MULTI / STRATA_ONE_TOKEN_COMMIT are hard-off; as opt-ins head-mix is exact and +0.4% decode on R9700 (patch inside)
- #1478 gfx1201 (R9700): the gfx1151 prompt switches in arch_defaults.cpp are exact and +4-8% prompt; suggest a gfx12 entry (decode half is slower)
- #1475 expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)
- #1474 hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
- #1469 0.1.40.2: decode on a Tesla P40 (sm_61) is about half of 0.1.40's; bisected to 2e4ddf6, and restoring __restrict__ in pdl.hpp (STRATA_PDL_RESTRICT) gives back most of it
- #1453 Live agent workload on RX 7900 XTX (24 GB) + 30 GB RAM — 70–74 tok/s, 91–96% expert cache hits
- #1445 vision encoder did not start on V100-SXM2-32GB
- #1431 Feature request: tool-call emission recovery for low-bit quants in agent workloads (Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS)
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buying
- #1403 Almost no changes from 2 recent updates,
- #1393 Support for AMD Radeon 890M APU
- #1353 stager-transient-regression
- #1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40
- #1347 serve: a full conversation cache drops the snapshot at the physical-RAM gate instead of evicting to pass it (270k tokens re-read, 96 s, 32 times)
- #1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
- #1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
- #1286 PCIe Issue with Strata and RTX 3060
- #1285 Dual RTX 5090: Strata 0.1.40 production decode at 243 tok/s and synthetic peer-prefill results
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1272 0.1.40: with the shared-expert stream fork off by default, decode on 2x RX 6900 XT (gfx1030, Linux) is 5-13% slower; STRATA_SH_STREAM=1 restores it, and the #884 stall does not come back with it
- #1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speed
- #1258 gfx1151 (Strix Halo) prompt: per-kernel profile of a 16K prompt (rocprofv3) - where would outside help be useful?
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
- #1250 [Linux][RTX 3080 12 GB, 62 GB RAM] 0.1.40.1 start OOMs the host while allocating the page-locked expert copy (IQ2_XS); 0.1.31 starts fine
- #1233 ``` Windows + Strix Halo (gfx1151, Radeon 8060S): UD-IQ4_XS runs end-to-end — decode flat 29.5-30.4 tok/s at 1K-32K, but a GPU driver timeout (TDR) during a cold 32K prefill ```
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???
- #1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores it
- #1168 Client-supplied stop sequences are ignored (OpenAI `stop` and Anthropic `stop_sequences`) — blocks agent/harness use
- #1155 [Windows / AMD] --vision cpu is still unreachable: hip_vision()'s WIN gate, find_vcvars()'s VS-18 range, and no encoder in the ready-made engine
- #1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
- #1118 Batch slots: the engine stalls ("no progress for 60 s … reading the prompt (batched)") when a long prompt is read while another slot decodes; STRATA_BATCH_DECODE_SHARE=0 avoids it
- #1103 hip: gfx1030 (RX 6800 XT) intermittent verify timeouts (#267): ROCm 7.14 gfx103X wheels regression, fixed by 7.13
- #1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40
- #1071 vmm.cpp: the CUDA 12 engine does not build with a toolkit older than 12.5 (missing CUDART_VERSION guard)
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
- #1059 serve: a request the engine rejects with ERR waits the full 300 s before failing (STOP + _drain_control("DONE") after an ERR), holding the control lock
- #1058 0.1.40: a tool call quoted in the thinking is delivered as a real call (fires on a max_tokens cut and inside code fences)
- #1056 Prefill: Stager threads yield-spin while waiting. On Linux with --mmap-experts (DGX Spark / GB10), prompts are read ~10x slower; elsewhere ~20 cores are burned during prefill
- #1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"
- #1008 Title: 0.1.39 decode is ~15–23% slower than 0.1.38 on AMD (gfx1201/HIP, 2× R9700); prefill unchanged
- #997 0.1.39: engine exits (code 1) with "parallel": 4 on UD-IQ4_XS, RTX 3090 24 GB, three concurrent 8K prompts (reproducible; IQ3_S unaffected)
- #990 MCP install planning rejects supported AMD CPU vision (and retains the old Windows HIP restriction)
- #980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missing
- #979 HIP: a gfx1201 hipBLASLt table for 1.2.1 (100201, ROCm 7.2.0) makes the prompt 6-17% slower than no table
- #977 --check prints [ok] for below-floor RAM and a misleading 'This PC can run Strata' verdict
- #976 serve: structured-output failure path returns error: null instead of 502 structured_output_failed (test_responses)
- #975 setup.get_prebuilt_hip() zip-unpack branch returns None on Linux (tools/test_setup_amd.py)
- #974 Missing --kv-resident in generated args when --resident-budget-gib is set (tools/test_setup_unsloth.py)
- #973 Golden drift: all 46 test_enter_for_every_question subtests fail on main (tools/test_setup_golden.py)
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
- #968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
- #959 ./setup.sh fails on Gentoo Linux
- #954 [BUG] 0.1.39: fused int8 prompt experts (STRATA_PF_FUSED=1) + #583 byte-budget ring → kernel hang, `prefill mmq: iota: unknown error` on sm_120 (RTX 5090, WSL2)
- #952 BrokePipeError when terminating Strata on Linux
- #942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39
- #918 Windows: Radeon 8060S (gfx1151) builds, selftests and passes the HIP ctest on top of #895
- #893 serve/server.py reads POST bodies as empty when sent with Transfer-Encoding: chunked (no Content-Length fallback) — breaks requests through some reverse-proxy/relay agents
- #892 Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works
- #884 HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)
- #881 [Windows / AMD] RX 6800M (gfx1031) runs with a self-built HIP engine — plus 2 Windows bugs found (packaging encoding, vision `vcvars`)
- #843 Tool calls silently dropped in agent sessions after an empty assistant turn (fix: skip empty assistant turns when rendering)
- #836 Milestone: Add Apple Silicon Cross-Platform Support
- #798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`
- #771 Linux expert arena (0.1.39): MADV_HUGEPAGE + defrag=madvise makes the arena's faults 20x slower (25 s start -> 432 s), and 'loaded ... at X GiB/s' does not show it
- #747 Feature request: multi-conversation parking on multi-GPU layer-split
- #710 v0.1.38: cross-request tool loop repeats an ineffective repair 98 times despite a thinking budget
- #703 [AMD HIP] RX 6750 XT (gfx1031) works in heterogeneous layer split; long multimodal prompts can crash the engine
- #702 [AMD HIP] RX 9060 XT (gfx1200) Repeating output
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
- #647 Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS works on double Vega 20 GPU (Radeon Pro VII 16gb each).
- #633 The pinned expert arena is allocated without a host-RAM check: a container limit or a low MemAvailable ends in the OOM killer mid-load, not a message (the file tier already has the probe)
- #606 After a 36,689-token request at 155k, every later request answers one repeated token until the engine reloads
- #602 Measured numbers: Qwen3.8-Flash-Next on 2× RTX 3060 (Q2_0 vs IQ3_S), plus a local reproduction of #75
- #601 Linux build fails on Ubuntu 26.04 (glibc 2.43) with CUDA 12.9 (nvcc host-math conflict)
- #597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)
- #566 Proposal: extend existing calibration to Linux HIP
- #560 AMD/Linux desktop: default VRAM reserve (700 MiB) lets the driver evict ~24 GB to RAM, OOM kills kwin; --vram-reserve-mib 3072 fixes it
- #549 Update.sh fails to update with a KeyError: 'args' error
- #541 AMD/HIP (R9700, engine 0.1.36): batched prompt read stalls at a repeatable token; STRATA_PLE_BATCH=0 avoids it
- #533 Hot VRAM resize: let the engine yield VRAM to other apps and take it back
- #506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)
Pull requests
- #1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegister
- #1492 fix(serve): save and restore sessions without an MTP drafter
- #1490 Add a workflow that builds and pushes the Docker image on a version tag
- #1489 Responses persistence and disk-only cache: CUDA/HIP integration
- #1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)
- #1471 feat: optionally offload idle session state while keeping model weights loaded
- #1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64K
- #1459 bench: RX 7900 XTX + 6800 XT helper layout, 1M and long-output results
- #1458 helper cache: configurable VRAM allowance with allocation-floor checks
- #1436 fix(serve): isolate CPU cores for independent Linux replicas
- #1434 telemetry: GPU load, VRAM, temperature and power for AMD cards on Windows (#1380)
- #1430 serve: a tool call written <function= NAME> is the call NAME
- #1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262K
- #1420 Add optional GLM-5.3-Flash support from Project Maya
- #1416 prefill: the CPU share's activations quantized on the GPU, copied down beside the routing sync (builds on #1414)
- #1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer split
- #1408 docs: a short "Contributing a change or report" section in AGENTS.md and a test-requests list
- #1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweep
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1399 qsa_select_bench: accuracy floor per sample (fails on an RTX 5090 / MSVC build)
- #1398 docs: keeping Codex CLI off the network with Strata
- #1396 hip/gfx906: the tree does not build (q6_k MMQ instance and cudaEventBlockingSync)
- #1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
- #1383 generate: say that CUDA_LAUNCH_BLOCKING=1 hangs the verify windows (#1341 #964)
- #1378 bench: community report, RX 6600 8 GB (gfx1032), EPYC 7232P, IQ3_XXS 128K on Linux
- #1377 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.2
- #1362 serve: preserve Markdown state across skipped tool-call newlines
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)
- #1359 setup: link the engine against the toolkit whose nvcc builds it (CUDAToolkit_ROOT)
- #1358 build: vmm.cpp builds with a CUDA toolkit older than 12.5 again (#1071)
- #1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweep
- #1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)
- #1323 prefill: batched unbuffered stager reads; close the experts.bin view once reads are unbuffered (port of #833)
- #1318 docs(AMD_HIP): my R9700 numbers for hipBLASLt 1.4.1 and STRATA_HIP_WMMA
- #1290 Fix shared MTP slot initialization and speculative row bounds
- #1289 hip: the gfx1100 100401 table covers the small-T expert GEMM (prompt +31.9%)
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
- #1270 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.40.1 - one card, layer split and expert helper; stock, the two gfx103x switches, and #1150 #1151 #1167
- #1260 tests: order MMQ parity transfers on the compute stream
- #1227 mcp: the install plan follows setup's AMD rules - Windows AMD and vision=cpu on Linux are planned (#990)
- #1226 Community benchmark: 2x RTX PRO 4500 Blackwell, Swift IQ3_XXS, engine 0.1.40.1 (solo to 256k, --batch after #776)
- #1224 serve: the Monitor's AMD GPU readings on Windows through ADLX (load, VRAM, temperature, power)
- #1202 dashboard: Serve the web page under an optional dashboard path
- #1185 verify: a window graph that finds no VRAM frees older batch slot layouts' graphs instead of failing (#997)
- #1183 serve: release batch control after terminal native ERR
- #1166 cpu pool: --host-core sibling, the host off the interrupts' logical processor (hybrid CPUs too)
- #1156 setup: reach the CPU image encoder on Windows + AMD (#1155, #881)
- #1136 Add opt-in durable chat history, legacy import and browser compaction
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
- #1125 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)
- #1122 Pipelined windows + async adaptive tier together (--pipeline-windows 2 --adapt-async 1)
- #1120 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, opt-in)
- #1105 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vision footprint
- #1101 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)
- #1092 launcher: pick a model and its settings in one window (experimental)
- #1090 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #1087 cpu: allow configuring expert-pool tasks per phase
- #1086 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
- #1057 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)
- #1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)
- #1039 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #1017 feature: RDNA3 support / Cache-aware routing
- #1016 Bench: RTX 5060 Ti 16 GB on Ryzen 9 7945HX, IQ3_S
- #1013 serve: a continued batch request reports its own prompt's reuse; input_tokens never 0
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)
- #1003 resident RAM mode: release each layer's mapped experts as it is copied (Windows)
- #1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)
- #1001 generate: wait for the residency table's upload before the next window (#871)
- #987 serve: /props reports chat_template_caps so Zed offers tools (#986)
- #970 serve: a tool call written inside the thinking ends it implicitly (#804)
- #965 serve: pin Claude Code's per-request billing stamp, so the conversation cache reuses agent prompts
- #960 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #949 cpu: allow configuring expert-pool tasks per phase
- #946 launcher: pick a model and its settings in one window (experimental)
- #944 hip: gfx11 WMMA prompt attention and selection scorer, gfx1151 hipBLASLt table
- #943 Add experimental deepMoE Vulkan backend for DeepSeek text chat
- #936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…
- #904 Verify window: PDL on sm_90+, graph branches, batched Q4_0 KV append and PLE post-ops (bit-identical)
- #902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)
- #897 setup: the disk check counts only what step 6 still writes
- #895 hip: support Radeon gfx1151 APUs
- #889 sycl: read doorbell flags with an uncached L1+L3 hint (the GPU never saw the host's store on an Arc Pro B60)
- #878 # Support sm_60, sm_70, and sm_86 with CUDA 12.6 or 12.8
- #869 serve: opt-in recovery from repeated reasoning passages
- #860 Unsloth UD-Q6_K_XL: Q8_0 PLE rows and Q6_K gate/up experts
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
- #851 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)
- #839 Add experimental Linux gfx900 setup with wave64 validation
- #833 Prefill: batched unbuffered expert reads, close the experts.bin view; opt-in RAM budget with helper caches
- #812 serve: Codex MCP tools (web search) fail as 'unsupported call': map namespace__name calls back to their namespace
- #811 Add an RTX 5090 + Ryzen 7 9700X IQ3_S benchmark on engine 0.1.39
- #810 serve: accept json_schema roots that are unions of object schemas
- #807 cuda: opt-in batched expert uploads to reduce host submission overhead
- #802 Run IQ1_S (Unsloth UD-IQ1_S) on the GPU and AVX-2, and read the expert file tier faster
- #797 sycl: enable Arc A770 inference and document measured performance
- #796 generate: one startup VRAM plan before the expert cache is committed (#765)
- #792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layer
- #789 prefill: overlap shared-expert work with routed expert uploads
- #788 serve: tolerate broken stdin pipes during engine cleanup
- #787 fix: parse tool calls inside reasoning
- #784 sycl: follow #626 (ThreadAffinity) and #559 (NativeDense::load layer range) so v0.1.39 compiles
- #777 Add an RTX 5090 + Ryzen 7 9700X IQ3_XXS benchmark on engine 0.1.39
- #773 file tier: unbuffered reads on Linux too (O_DIRECT + the kernel's asynchronous reads)
- #750 docs: describe a targeted ROCr USERPTR reclaim workaround
- #734 Reuse long chat history when the last user message is edited
- #733 Pipeline CPU expert work without a batch-wide phase barrier
- #732 Overlap streamed KV uploads with prefill computation
- #731 Avoid duplicate expert slots across helper and primary GPUs
- #724 Add local PDF, Word and Excel attachments with screenshot previews
- #718 serve: park conversations on disk, keyed by a client id; requests without one stay in RAM
- #717 fix: make the CUDA language link the shared runtime (Windows sm_70 LNK1169, #585)
- #716 conversation parking with --layer-split
- #715 serve: say Connection: close on every response
- #712 Fix Windows AMD telemetry and add optional Arabic UI
- #711 kv: --kv k8v4 streams with --kv-resident
- #700 serve: a <tool_call> named in prose before the real call is content, not a malformed call
- #699 Linux: read ahead at startup (cold start ~920 s → 70 s)
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)
- #695 hip: the tuning table's arch gate accepts the RDNA4 sibling (gfx1200 <-> gfx1201)
- #688 Bench: RTX 5090 / Ryzen 9 9950X3D, IQ3_S, v0.1.34 vs v0.1.38 vs fused prompt kernels
- #686 Add a tiny zero-dependency terminal launcher for Strata
- #685 Add optional Windows HIP vision encoder builds and package validation
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #677 bench: Linux + 2x AMD Instinct MI50 16 GB (gfx906), Coder IQ1_M, 4K to 128K prompt tokens
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)
- #660 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)
- #655 sm_75: run the prompt path's BF16 products on the FP16 tensor cores (+15-18 % prefill on RTX 2080 Ti)
- #650 pinned: back the Linux expert arena with transparent huge pages
- #640 STRATA_ARENA_MMAP=1: the expert arena as a read-only mapped file (Linux), for small-RAM machines
- #637 serve: retry an engine start that fails, and keep the last known context after a failed restart
- #636 native: add opt-in GBNF constraints to existing generation
- #635 serve: a malformed "tools" value is a 400 naming the field, not a closed connection (#592)
- #634 strata_pack.py build: refuse an output directory that already holds a pack, naming the file
- #632 expert profile: refuse a header version other than 1, naming both numbers and the file
- #626 cpu: preserve full thread affinity on Windows and Linux
- #618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokens
- #611 cpu: group ARM capacities and honor the pool's reserved host core
- #594 serve: read the request body an answer left unread, so the close is not a reset
- #570 setup: a read-only --gguf-dir no longer stops setup when marking checked shards
- #569 serve, setup: an empty API key is refused from the config file and from setup's --api-key (#213)
- #567 serve: tokenise the prompt from the last shared prefix, not from the start
- #563 Give the expert cache's VRAM back without unloading the model (#533)
- #508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
- #491 docker: GGUF_DIR, RESIDENT_BUDGET_GIB and KV_STREAMING env vars; document the stop timeout
- #488 pinned: honor STRATA_NO_LARGEPAGES on Linux; name the hugetlb pool shortfall
- #482 Networked Pool support
- #476 native head: report the cudaHostAlloc error, not just "cannot pin N MiB"
- #472 tests: test_prebuilt_hip_zip passes on Linux (EXE as on Windows)
- #470 setup: a download stopped before its rename is finished without a request (no 5-minute 416 retry loop)
- #456 Monitor: the conversation cache card (parked conversations from the engine's CACHE lines)
- #453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-resident
- #452 qsa_prompt_attn: Q4_0 KV on tensor cores (mode 4) - long prompts with --kv q4_0 ~19% faster
- #439 prefill: a group's expert gathers in one launch (bit-identical, long prompts +4%)
- #435 setup: resuming a download asks only for the missing space, and a finished .part is not re-requested (#425)
- #428 tests: test_setup_golden passes on Linux (<EXE> only as a whole path part)
- #427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)
- #426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)
- #413 prefill: DeltaNet recurrence with the three value heads of a key head in one thread (bitwise; 1.4x on a 4080 SUPER, 1.3x on a 3090)
- #398 setup: stop on EOF instead of accepting prompt defaults
- #389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262K
- #382 HIP: MTP prompt pass per group by default (fixes prompt hang on gfx1201 with KV streaming)
- #381 Add HIP image encoder support and Linux AMD GPU telemetry
- #380 HIP: count the desktop's VRAM on Windows (WDDM budget, STRATA_WDDM_BUDGET)
- #377 HIP: time the PCIe probe on the host on Windows
- #376 HIP: a failed first configure no longer leaves the kernels at -O0
- #373 HIP: build with Visual Studio 2026's C++ library (MSVC 14.51+)
- #362 File tier: read the GGUF in place unbuffered when the file cache cannot keep it beside the RAM budget (Windows; #286 ported into the existing tier)
- #360 qsa_prompt_attn: Volta (sm_70) crashes in the prompt path - compare the full compute capability
- #358 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)
- #357 Start: read the expert arena unbuffered when the file cache cannot help (Windows; part 1 of #285)
- #356 HIP: build the AMD engine on Windows
- #353 NVFP4 routed experts on 0.1.35 (experimental): 120a only for the W4A4 unit (rebase of #292)
- #325 AMD working on Windows (9070XT tested)
- #324 fix(setup): install the engine, model files and packages the checkout was tested with
- #323 hip: add experimental RX 5500 XT 8 GB (gfx1012) support
- #302 HIP/Windows: batch expert copy dependencies in long prefill
- #294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windows
- #286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.bin
- #275 Add bounded disk persistence for evicted conversations
- #272 feat(cpu): add hybrid architecture awareness and --pool-affinity for P/E-core CPUs
- #247 HIP: Windows and gfx1201 (RX 9070)
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)
- #202 ple: --ple-io ram locks the n-gram table in RAM (no SSD read on the prompt/token path)
- #190 Spill evicted conversation snapshots to a bounded optional disk cache
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)
- #158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUs
- #154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)
- #139 perf(prefill): add shape-selected pre-Ampere BF16 SGEMM with current-main A/B
- #136 Linux prebuilt from CI: CUDA 12.8, sm_80/86/89 + PTX, attached to each release
- #130 setup: Volta (V100, sm_70) and Turing as an experimental build path
- #129 core: add optional shared expert arena backing
- #110 Multi-GPU: a second GPU as an expert tier for decode and prompts (--peer-device), + multi-conversation cache
- #107 serve: report cached prompt tokens, llama.cpp timings and GET /v1/status
- #105 The Monitor's GPU readings on ROCm: an amdgpu sysfs backend, and the disk rate without psutil
- #104 The Monitor's live tok/s is a rate over the last seconds, not the mean since the first token
- #101 build: allow sm_70 (Volta) builds — the only sm_80+ dependency is unreferenced
- #94 Add experimental gfx1100 HIP backend for RX 7900 XTX
- #92 setup: name a missing or short shard with its numbers; verify reads back the manifest sha256
- #91 GGUF reader: refuse a duplicate tensor name at open, naming the tensor and the file
- #89 loader: bulk-read the expert arena and the MTP drafter (MSVC splits ifstream reads into 4095-byte calls)
- #88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)
- #87 Turing port: run on RTX 20 (sm_75) and newer
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)
- #79 linux: make DirectFile reads actually asynchronous
- #67 Support OrcaRouter IQ3_XXS with explicit BF16 compatibility packing
- #52 nvme kv cache
- #41 serve: preserve primary conversation state across auxiliary calls
- #22 tools: add real-time web dashboard for hardware and engine monitoring
- #14 CPU pool: fix dangling else in Linux physical_cores() (fixes #13, root cause of #11)
- #9 CPU pool: sleep between requests instead of spinning (fixes #4)