Labels / Server & API
Server & API
Auto topic: Server & API
Issues
- #1487 Windows/CUDA 11.8 Volta integration: focused patches and reproducible handoff
- #1484 [SYCL] 2x Arc Pro B70 + UD-IQ4_XS: the port's default expert kernels give corrupted output (silent wrong answers); STRATA_EXPERT_SPLIT=1 restores it
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
- #1468 prefill copy_i32: illegal memory access on long prompts when the prefill chunk is large (auto:32768)
- #1463 MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)
- #1453 Live agent workload on RX 7900 XTX (24 GB) + 30 GB RAM — 70–74 tok/s, 91–96% expert cache hits
- #1445 vision encoder did not start on V100-SXM2-32GB
- #1444 can kvcache be saved and reused on disk?
- #1442 Feature request: support chat title generating from Xcode's agent
- #1440 [Bug]: [SYCL] 0.1.40.2 on 2x Arc Pro B70: --layer-split fails three ways (#1054's fix from #1111 is not in main, plus two new split regressions)
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdog
- #1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per request
- #1392 Web app: an assistant reply with no content is dropped from the conversation history (reasoning models repeat themselves)
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by default
- #1385 Intermittent 'int' object has no attribute 'ranks' in tokenizer
- #1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of it
- #1364 Metrics logs & live dashboard
- #1353 stager-transient-regression
- #1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)
- #1347 serve: a full conversation cache drops the snapshot at the physical-RAM gate instead of evicting to pass it (270k tokens re-read, 96 s, 32 times)
- #1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the model
- #1322 setup re-run with --vision gpu does not add --vision to an existing config's args
- #1317 serve: an untimed read of the engine/vision pipe can hold the request turn forever (intermittent hang; the #481 lost-step class)
- #1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)
- #1302 0.1.40 SYCL port on 2x Arc Pro B65: expert-cache fill stalls (1-core spin + xe "Timedout job" + device coredump); 0.1.38-based engine works on the same card
- #1286 PCIe Issue with Strata and RTX 3060
- #1285 Dual RTX 5090: Strata 0.1.40 production decode at 243 tok/s and synthetic peer-prefill results
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)
- #1275 Windows HIP: --prefill auto can pick a chunk whose VRAM borrow leaves too little for the verify captures — silent engine exit at serve init
- #1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speed
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert pool
- #1250 [Linux][RTX 3080 12 GB, 62 GB RAM] 0.1.40.1 start OOMs the host while allocating the page-locked expert copy (IQ2_XS); 0.1.31 starts fine
- #1234 Windows/RTX 4090 (0.1.39): stall report stage "decode -1" for 61 s; 0 layers served - thread stacks attached
- #1230 serve/web: add English / Simplified Chinese (zh-CN) UI localization
- #1229 serve: phone photos reach the model rotated by 90° (EXIF Orientation is ignored)
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patch
- #1206 DirectStorage implementation
- #1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???
- #1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores it
- #1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)
- #1180 STRATA_PF_FUSED on gfx1100: missing __syncthreads() in native_w11_kernel + waves_per_eu(8) wrong on clang 22 / ROCm 7.10
- #1178 serve: Special tokens (<|im_end|>) emitted during model thinking cause tasks to terminate early
- #1177 Feature request: add timestamps to engine log lines
- #1169 k8v4 + --kv-resident: the batched KV append (#783) skips the host mirror for hybrid — parking and KV streaming serve stale bytes (cross-conversation contamination on 0.1.40)
- #1168 Client-supplied stop sequences are ignored (OpenAI `stop` and Anthropic `stop_sequences`) — blocks agent/harness use
- #1155 [Windows / AMD] --vision cpu is still unreachable: hip_vision()'s WIN gate, find_vcvars()'s VS-18 range, and no encoder in the ready-made engine
- #1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all pass
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)
- #1128 0.1.40 --kv-grow: a short request trims the K/V to 8192 cells and drops the conversation cache, so the next 128K+ continuation is re-read in full (20.9% of continuations, RTX 5090, IQ3_S)
- #1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML type
- #1118 Batch slots: the engine stalls ("no progress for 60 s … reading the prompt (batched)") when a long prompt is read while another slot decodes; STRATA_BATCH_DECODE_SHARE=0 avoids it
- #1113 sycl: the v0.1.40.1 port does not build as shipped (3 compile errors, 12 undefined symbols)
- #1094 --layer-split auto fails on dual GPUs for 512k context
- #1072 Vision: >64 unique images can evict .sve files still referenced by the current request
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)
- #1059 serve: a request the engine rejects with ERR waits the full 300 s before failing (STOP + _drain_control("DONE") after an ERR), holding the control lock
- #1058 0.1.40: a tool call quoted in the thinking is delivered as a real call (fires on a max_tokens cut and inside code fences)
- #1056 Prefill: Stager threads yield-spin while waiting. On Linux with --mmap-experts (DGX Spark / GB10), prompts are read ~10x slower; elsewhere ~20 cores are burned during prefill
- #1053 A reply that ends inside `<think>` (no `</think>`) comes back empty with `finish_reason: "stop"`, and the agent stops — close the thinking and continue once
- #1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)
- #1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
- #1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"
- #1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)
- #997 0.1.39: engine exits (code 1) with "parallel": 4 on UD-IQ4_XS, RTX 3090 24 GB, three concurrent 8K prompts (reproducible; IQ3_S unaffected)
- #990 MCP install planning rejects supported AMD CPU vision (and retains the old Windows HIP restriction)
- #986 Zed says Tools Unsupported: `/props` has no `chat_template_caps`
- #980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missing
- #978 Strata 7900xtx Qwen 3.8 Flash Next Q3_xxs VS Openai gpt6 Luna
- #976 serve: structured-output failure path returns error: null instead of 502 structured_output_failed (test_responses)
- #973 Golden drift: all 46 test_enter_for_every_question subtests fail on main (tools/test_setup_golden.py)
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)
- #968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctly
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_S
- #954 [BUG] 0.1.39: fused int8 prompt experts (STRATA_PF_FUSED=1) + #583 byte-budget ring → kernel hang, `prefill mmq: iota: unknown error` on sm_120 (RTX 5090, WSL2)
- #953 Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decode
- #952 BrokePipeError when terminating Strata on Linux
- #951 crash after first inputs
- #948 A40 - failed to run
- #942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39
- #941 UPD: Successfully run on RTX 3090 24gb + Z590E + 64gb ddr4 3200 2ch
- #937 `[gfx1030] HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION in gdn_step_commit_kernel after a few long fresh prefills`
- #929 `--batch` with `--layer-split` exits ("verify batch: layer 34 never rang") with a 3070 first and a P40 second; works with the P40 first (CUDA, v0.1.39)
- #925 Setup: show progress in the browser from the first second (status page with steps, speed and ETA)
- #922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runs
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
- #915 Windows AMD report: RX 6800M (gfx1031, laptop) works with a self-built HIP engine — please add 0x73DF / gfx1031 to the Windows path
- #914 strata-vision helper leaks /tmp/strata-vision-* temp dirs (unbounded growth -> OOM on tmpfs /tmp)
- #909 Proposal (docs only): a short "contributing a change or report" section for AGENTS.md, and a pinned list of test requests. Yes/no?
- #893 serve/server.py reads POST bodies as empty when sent with Transfer-Encoding: chunked (no Content-Length fallback) — breaks requests through some reverse-proxy/relay agents
- #884 HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)
- #879 All-logits-NaN degeneration (one repeated token forever) still fires on 0.1.39b / current main — first poison traced to the MoE output rows: garbage bits, varying layer
- #875 P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)
- #871 Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)
- #870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current main
- #867 Intel SYCL port on 2x Arc Pro B60: results, and "never rang (graph finished)" on the host-mirror path
- #843 Tool calls silently dropped in agent sessions after an empty assistant turn (fix: skip empty assistant turns when rendering)
- #841 not stable start with 3+ cards
- #840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaningless
- #831 The profile's length caps the expert arena — even an explicit --expert-cache N
- #830 Let a one-shot request skip the prompt-cache work (~45 ms per call that is never reused)
- #828 Heap corruption (munmap_chunk: invalid pointer) at startup in Prefill::init — crashes with 1/2/4 GPUs
- #817 Assistant turn can be interrupted when generation emits a reserved ChatML boundary marker
- #804 Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`
- #803 Perplexity ~7-10% above llama.cpp on the same GGUF (Qwen3.8-Flash-Next), also at 2K context, every version tested
- #795 Main engine strata.exe crashes with STATUS_ILLEGAL_INSTRUCTION (0xC000001D) on non-AVX-512 CPU - 20 WER crashes across v0.1.35/v0.1.38, fault offsets cluster in fixed code regions
- #791 Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split
- #779 laptop crash with latest
- #772 Community project: strata-router, a multi-node router for Strata / 社区项目:Strata 多节点路由入口
- #771 Linux expert arena (0.1.39): MADV_HUGEPAGE + defrag=madvise makes the arena's faults 20x slower (25 s start -> 432 s), and 'loaded ... at X GiB/s' does not show it
- #769 server.py: error: port 8080 is already in use
- #756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)
- #754 Anthropic API: tool calls intermittently emitted inside `thinking` and returned as `end_turn`
- #747 Feature request: multi-conversation parking on multi-GPU layer-split
- #738 Configurable system prompt in the web GUI
- #710 v0.1.38: cross-request tool loop repeats an ineffective repair 98 times despite a thinking budget
- #691 Windows: generation slows when the server window is minimized (EcoQoS moves threads to E-cores) — fix included
- #690 Layer split: native prefill kernels fail on the second GPU (no kernel image available), regardless of which card it is
- #675 Feature request: `logprobs` / `top_logprobs` on `/v1/chat/completions` (typed decisions from one forward pass)
- #671 serve: ordered lists whose items are separated by blank lines render every item as "1."
- #670 Keeping the old engine when an update succeeds but turns out bad - and a question about where the web app should call it
- #669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
- #654 HIP (gfx1200) RuntimeError Exception Code: 0xC0000005
- #644 Specifying explicit GPU split crashes Strata
- #629 Re-running setup with a different --context silently discards mcp_servers, mcp and sampling from the run config
- #625 Vision (CPU path): offer a Q8_0 mmproj and a higher image-token cap
- #620 IQ3_S: "native head upload: out of memory" at startup when the desktop runs on a second GPU (more free VRAM on the card)
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same box
- #616 RTX A3000 / 0.1.38: AVX2 i-quant gate/up codebook gather — bit-exact ~1.3x isolated, +1.0% e2e, kept as opt-in
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #610 RTX A3000 / 0.1.38: a native IQ3_XXS decode window decomposes via STRATA_VERIFY_PROFILE, and its GPU side is memory-bandwidth-bound
- #606 After a 36,689-token request at 155k, every later request answers one repeated token until the engine reloads
- #605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)
- #604 Resident expert placement on a layer split, for mixed or older GPUs: where I think it helps, plus some doc notes, and an offer to test
- #596 Add Conversation Cache Management UI to Monitor Tab
- #592 serve: a malformed "tools" value kills the request thread (connection reset / 502 instead of a 400)
- #590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 W
- #588 Reported expert-cache hit rate excludes --pcie-frac experts, so it rises as tok/s falls
- #579 AMD/HIP RX 9070 XT (gfx1201) stalls during batched prompt ingestion on ROCm 10.0
- #560 AMD/Linux desktop: default VRAM reserve (700 MiB) lets the driver evict ~24 GB to RAM, OOM kills kwin; --vram-reserve-mib 3072 fixes it
- #556 EXL3 (exllamav3 trellis) weight support
- #548 decode_cluster_parity: graph replays race their input upload (pageable cudaMemcpy + non-blocking stream)
- #541 AMD/HIP (R9700, engine 0.1.36): batched prompt read stalls at a repeatable token; STRATA_PLE_BATCH=0 avoids it
- #537 Literal </think> quoted in reasoning yields empty stop or leaked reasoning (v0.1.37)
- #535 Thinking speed suddenly drops to 0.5 token/s
- #528 Conversation cache: decode drops ~4-6x on continued conversations (0.1.36, Windows, RTX 5090, IQ3_XXS)
- #519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt path
- #517 new error `raise RuntimeError("the engine exited before it was ready" + (f" (see {log})" if log else "") +`
- #516 desktop (KDE/Wayland): compositor VRAM exhaustion when expert-cache auto fills a 24 GB card - pin framebuffer -12, KWin graphics reset
- #511 Engine stopped unexpectedly
- #506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversed
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)
Pull requests
- #1492 fix(serve): save and restore sessions without an MTP drafter
- #1489 Responses persistence and disk-only cache: CUDA/HIP integration
- #1486 feat: introduce Dockerfile.rocm-stable
- #1481 web: add Copy below code blocks (#1257)
- #1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)
- #1476 Update GLM port to Project Maya v1.0.4
- #1471 feat: optionally offload idle session state while keeping model weights loaded
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)
- #1455 serve: reuse prompt tokens across safe shared boundaries
- #1450 web: wide layout — a four-column single page on very wide screens, and a card look for the chat and the drawer
- #1449 Add Metal backend support for Strata on macOS
- #1448 #606 follow-up: clamp the three remaining unclamped q8_1 scale/sum emit sites
- #1436 fix(serve): isolate CPU cores for independent Linux replicas
- #1434 telemetry: GPU load, VRAM, temperature and power for AMD cards on Windows (#1380)
- #1432 bench: Arc Pro B65 Gen4 results on Strata v0.1.40.2
- #1430 serve: a tool call written <function= NAME> is the call NAME
- #1427 web: Russian UI localization (dictionary + header language picker)
- #1426 dflash: experimental standalone DeepSpec DFlash block drafter for Qwen3.8-Flash-Next (greedy, correctness-first)
- #1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)
- #1420 Add optional GLM-5.3-Flash support from Project Maya
- #1419 bench: replay token IDs and verified windows in serve decode
- #1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer split
- #1405 tools: add configurable launches and reuse local model downloads
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cards
- #1398 docs: keeping Codex CLI off the network with Strata
- #1396 hip/gfx906: the tree does not build (q6_k MMQ instance and cudaEventBlockingSync)
- #1395 hip: a hipBLASLt tuning table for gfx1150 (Radeon 890M), and gfx1151's exact speed switches on gfx1150
- #1394 serve: evict parked conversations to pass the physical-RAM gate instead of dropping the snapshot
- #1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixes
- #1383 generate: say that CUDA_LAUNCH_BLOCKING=1 hangs the verify windows (#1341 #964)
- #1379 prefill: STRATA_PREFILL_CPU_SHARE=auto times layers with and without the share and shares only while that is faster (#1282)
- #1374 experts: IQ1_S on GPU/AVX-2, IQ1_S+IQ1_M on AVX-512, and a measured PCIe share
- #1365 Dashboard - metrics logs
- #1362 serve: preserve Markdown state across skipped tool-call newlines
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)
- #1331 Aporte/disk mirror y sesiones agnosticas
- #1328 serve + engine: constrained decoding for JSON formats; accept tools together with json_schema
- #1327 serve: opt-in finite video input
- #1326 serve: support IPv6 bind addresses
- #1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)
- #1319 serve: add opt-in bounded Responses history and continuation
- #1314 serve: keep unfinished tool calls visible as content
- #1309 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
- #1290 Fix shared MTP slot initialization and speculative row bounds
- #1288 serve: a long read gives way by what is left to read, not prompt length (#656)
- #1282 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Halo
- #1274 setup: --thinking / --instruct set the sampling every client gets (#1129)
- #1273 serve: session files work with --peer-device
- #1271 serve: --conversation-cache-spill-dir keeps evicted conversations on disk across restarts
- #1269 serve: restore a session file without holding its K/V in RAM
- #1266 Add the Update the engine card
- #1265 Add the engine updater: verify, stage, all-or-nothing swap, rollback
- #1262 STRATA_HC_REQ8: hyper-connection projections requantized to int8 + fp32 scale per 32 at load (rebased #704)
- #1255 serve: /props total_slots is the engine's batch slots
- #1251 --peer-device prompt share without P2P: host route, per-token sums, FP16 transfers (2.1-2.5x prompts on a no-P2P pair)
- #1249 Pipelined batch: two requests run at half speed (pad rows in group windows, unbalanced slot choice)
- #1241 serve: the Responses API reads Codex's additional_tools input items as tools (#782)
- #1237 perf(expert): the cache fill's stage buffers pinned with cudaHostAlloc
- #1231 kv-grow: a run that cannot lend cache slots maps the whole window up front instead of writing past its first 16K cells
- #1228 serve: list items separated by blank lines are one list, and a list continued later keeps its numbers (#671)
- #1227 mcp: the install plan follows setup's AMD rules - Windows AMD and vision=cpu on Linux are planned (#990)
- #1224 serve: the Monitor's AMD GPU readings on Windows through ADLX (load, VRAM, temperature, power)
- #1223 generate: allow the elastic K/V with --peer-device (VMM ranges keep their memory on the owning GPU)
- #1222 serve: render a Codex compaction with that conversation's tool prefix
- #1221 serve: answer Codex thread-title turns without reading them
- #1218 Check the ready-made engine against GitHub's SHA-256 before installing it
- #1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoring
- #1207 fix(serve): keep cached images alive through request preparation
- #1204 serve: a Chinese / English language switch for the web app
- #1202 dashboard: Serve the web page under an optional dashboard path
- #1196 serve: share the thinking budget with other apps, like max_tokens and…
- #1195 serve: a POST /settings body without "defaults" must not clear every …
- #1190 resident RAM mode with a layer split: the copy keeps each stage's lend region
- #1189 Add assistant prefill support for chat completions
- #1187 gfx906: opt-in Q2_0 gate/up weight reuse without singleton regression
- #1183 serve: release batch control after terminal native ERR
- #1182 serve: poll cancellation during quiet prefill waits
- #1179 batch: with --batch-groups, a group's lowest free slot first (+14-22% at 2-3 clients)
- #1175 serve: a tool's "parameters" / "input_schema" that is not an object is a 400 naming it, not an AttributeError on its first call (#592 follow-up)
- #1172 serve: opt-in recovery of tool calls the model writes next to the template's form ("tool_call_recovery": true)
- #1162 Verify the downloaded engine against GitHub's SHA-256 (replaces #645)
- #1158 bench: community reports, RTX 4090, IQ3_XXS 204800 and IQ3_S 143360 - 0.1.38 to 0.1.40.1
- #1156 setup: reach the CPU image encoder on Windows + AMD (#1155, #881)
- #1152 Fix Windows AMD telemetry and add optional Arabic UI
- #1144 serve: add opt-in instruction skills to Chat
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776
- #1136 Add opt-in durable chat history, legacy import and browser compaction
- #1134 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload (re-submission of #466)
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
- #1132 docs: running Strata behind llama-swap
- #1131 generate: the prompt loan keeps its residency bookkeeping without the token graph (#796 part B)
- #1126 Harden server, installer and engine against high-risk bugs
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)
- #1119 Add local PDF, Word and Excel attachments with screenshot previews
- #1117 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in; resubmission of #1030 on the new main)
- #1114 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1112 web: make MCP tool activity readable and collapsible
- #1109 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
- #1106 web: follow streaming Thinking and preserve reading position
- #1105 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vision footprint
- #1102 fix: batch host paths honor refreshed residency (deadlocks the engine during a prompt loan)
- #1101 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)
- #1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)
- #1093 Monitor: Load/Unload split button + idle-unload slider, load busy guard, honest unload
- #1092 launcher: pick a model and its settings in one window (experimental)
- #1090 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #1084 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
- #1075 web: follow streaming Thinking and preserve reading position
- #1068 serve: a tool call from the reasoning only when it ends the turn (#1058)
- #1067 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8
- #1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer split
- #1060 Monitor: per-GPU view, more cards, and a layout you can rearrange
- #1057 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)
- #1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)
- #1050 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)
- #1048 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installed
- #1046 serve: a request's combined image file is deleted when the request is refused or never runs
- #1045 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)
- #1044 serve: image sources - network paths refused, URLs capped and fetched outside the FIFO, no local files for other origins' pages
- #1043 generate: residency-table uploads wait for their own copy before the non-blocking streams read the table
- #1041 serve: an image in an Anthropic tool_result reaches the encoder (Claude Code's Read of a picture)
- #1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #1038 decode: STRATA_PCIE_BALANCE=1 picks each layer's PCIe count from measured costs (opt-in)
- #1035 serve: share the thinking budget with other apps, like max_tokens and…
- #1032 serve: add opt-in bounded Responses history and continuation
- #1031 serve: a POST /settings body without "defaults" must not clear every …
- #1030 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in, CUDA VMM)
- #1027 bench: community report, RX 6700 XT 12 GB (gfx1031), Ryzen 7 5700G, I…
- #1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)
- #1017 feature: RDNA3 support / Cache-aware routing
- #1014 vision: downscale large images before CPU encoding (12MP: >8 min → 20 s)
- #1013 serve: a continued batch request reports its own prompt's reuse; input_tokens never 0
- #1011 kv: one pinned KV pool shared by the sessions (--kv-pool-tokens; 4 lanes x 256K pin 6.24 GiB instead of 15.5)
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)
- #1004 serve: /props total_slots is the engine's batch slots
- #1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)
- #1001 generate: wait for the residency table's upload before the next window (#871)
- #998 serve: an engine that exits after an unread ERR line says why (#997 #890)
- #992 security-hardening
- #988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)
- #987 serve: /props reports chat_template_caps so Zed offers tools (#986)
- #970 serve: a tool call written inside the thinking ends it implicitly (#804)
- #969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)
- #966 serve: recover tool calls the model writes next to the template's form, and return the rest as content
- #965 serve: pin Claude Code's per-request billing stamp, so the conversation cache reuses agent prompts
- #963 feat(server): stop unchanged completed tool-call loops
- #960 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it again
- #956 serve: an API key with spaces around it or outside Latin-1 works, and a refused key is not called a missing one (#725)
- #946 launcher: pick a model and its settings in one window (experimental)
- #944 hip: gfx11 WMMA prompt attention and selection scorer, gfx1151 hipBLASLt table
- #943 Add experimental deepMoE Vulkan backend for DeepSeek text chat
- #939 serve: a tool call written inside an unclosed thinking span is an implicit think end (#804)
- #934 Per-layer expert cache: each layer's slots are that layer's blob
- #931 serve: control-token text inside a message stays text (<|im_end|> in a file no longer ends the turn)
- #928 Gyro (TQ2_T / TQK6 / TQK7 + Hadamard rotation): native support for the agentionai Qwen3.8-Flash-Next quants
- #927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854
- #924 serve: render Codex's compaction with the conversation's tool prefix (Codex compatibility)
- #923 serve: answer Codex's thread-title requests without running them (Codex compatibility)
- #920 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)
- #904 Verify window: PDL on sm_90+, graph branches, batched Q4_0 KV append and PLE post-ops (bit-identical)
- #901 Add ARM64 NVIDIA GB10 build and native-pack support
- #895 hip: support Radeon gfx1151 APUs
- #894 serve: decode chunked request bodies (fixes empty-body 400 behind some proxies)
- #886 serve: skip completed empty assistant turns in chat history
- #880 layer split: the stage weight trim works under `auto` too 4X gain in the prefill in 4 way ring
- #876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)
- #869 serve: opt-in recovery from repeated reasoning passages
- #861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
- #852 engine: --gpu selects the visible GPUs without environment variables
- #848 Resident RAM mode with a layer split, --vram-reserve-later-mib (#642)
- #846 Keep MTP drafting in concurrent batch slots
- #845 verify: batch windows honor the all-resident stage (no host doorbells)
- #839 Add experimental Linux gfx900 setup with wave64 validation
- #837 server: "return_progress": true puts the prompt's PP line on the stream as llama.cpp's prompt_progress
- #833 Prefill: batched unbuffered expert reads, close the experts.bin view; opt-in RAM budget with helper caches
- #829 docs: running Strata behind llama-swap
- #822 Add assistant prefill support for chat completions
- #819 frontend: images inside Anthropic tool_result blocks reach the model
- #815 bench: community results for an RX 6800 on Windows (0.1.39 against 0.1.38)
- #812 serve: Codex MCP tools (web search) fail as 'unsupported call': map namespace__name calls back to their namespace
- #810 serve: accept json_schema roots that are unions of object schemas
- #809 Intel arc 0.1.39 + perf fixes
- #807 cuda: opt-in batched expert uploads to reduce host submission overhead
- #805 Lazy vision encoder: per-GPU elastic expert caches
- #796 generate: one startup VRAM plan before the expert cache is committed (#765)
- #793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kit
- #792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layer
- #790 serve: honour OpenAI tool_choice "required" and a named function
- #788 serve: tolerate broken stdin pipes during engine cleanup
- #787 fix: parse tool calls inside reasoning
- #785 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installed
- #764 decode: STRATA_ADAPT_LAG - #463's reproducibility without its stall
- #763 serve: enforce OpenAI tool_choice and allow structured output with tools
- #762 serve: extract JSON from prose and code fences
- #759 Add native Responses API support for Codex clients
- #757 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)
- #753 serve: isolate concurrent request metrics (depends on #559)
- #748 serve: invalidate the engine after fatal verification timeouts
- #745 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly, serve-path decode (Qwen3.8-Flash-Next IQ3_S)
- #727 Add opt-in durable chat history, legacy import and browser compaction
- #726 Add opt-in live cache resizing and desktop resource presets
- #724 Add local PDF, Word and Excel attachments with screenshot previews
- #718 serve: park conversations on disk, keyed by a client id; requests without one stay in RAM
- #715 serve: say Connection: close on every response
- #712 Fix Windows AMD telemetry and add optional Arabic UI
- #708 serve: --mtp is optional
- #701 serve: a malformed "tools" is a 400 naming the field (#592)
- #700 serve: a <tool_call> named in prose before the real call is content, not a malformed call
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)
- #688 Bench: RTX 5090 / Ryzen 9 9950X3D, IQ3_S, v0.1.34 vs v0.1.38 vs fused prompt kernels
- #686 Add a tiny zero-dependency terminal launcher for Strata
- #685 Add optional Windows HIP vision encoder builds and package validation
- #683 setup: a re-run of setup carries over the hand-edited config blocks (#629)
- #682 serve: a Chinese / English language switch for the web app
- #681 serve: a malformed "tools" is a 400, not a dead request thread
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #666 setup, serve: the image encoder's device as its own role (--vision-device)
- #665 serve: no automatic --layer-split when the args use --peer-device
- #663 prefill: a layer split's idle card streams a share of a one-chunk prompt's experts
- #656 serve: cooperative prefill preemption at safe chunk boundaries (--prefill-preempt)
- #653 serve: park conversations with a layer split
- #652 serve: commit only the tokens a verify window hands out
- #650 pinned: back the Linux expert arena with transparent huge pages
- #645 Update the engine from the web app, cautiously
- #637 serve: retry an engine start that fails, and keep the last known context after a failed restart
- #636 native: add opt-in GBNF constraints to existing generation
- #635 serve: a malformed "tools" value is a 400 naming the field, not a closed connection (#592)
- #630 serve: add opt-in stateless Responses API
- #627 opt for sm_70(just test V100-16g*1 or 2)
- #622 iq: gather the AVX2 codebook from the grid table (STRATA_IQ256_GATHER)
- #615 fix(serve): correct cache usage and timings across reasoning continuations
- #614 serve: optionally checkpoint an existing chunk near the prompt tail
- #608 setup: a 200K context between the 128K rule and 256K (#406)
- #594 serve: read the request body an answer left unread, so the close is not a reset
- #584 Split placement search
- #583 Prefill chunk ring
- #582 Serve tool result images
- #581 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vi…
- #580 Weight carve 0138
- #572 serve: return unfinished tool call text as content
- #571 Feat/responses api 451
- #569 serve, setup: an empty API key is refused from the config file and from setup's --api-key (#213)
- #568 parity tests: ple_parity runs without the missing fixtures, gr_parity and ple_parity wait for their setup
- #567 serve: tokenise the prompt from the last shared prefix, not from the start
- #563 Give the expert cache's VRAM back without unloading the model (#533)
- #562 serve: /slots reports n_prompt_tokens, so a llama.cpp-style context m…
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weights
- #555 serve: a request's combined image file is deleted when the request is refused or never runs
- #554 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)
- #553 serve: image sources - network paths refused, URLs capped and fetched outside the FIFO, no local files for other origins' pages
- #550 generate: residency-table uploads wait for their own copy before the non-blocking streams read the table
- #536 decode_cluster_parity: the graph replays wait for their uploads (the test failed now and then)
- #532 tests: prefill_mmq_kquant_test waits for its uploads before the kernels
- #529 serve: an image in an Anthropic tool_result reaches the encoder (Claude Code's Read of a picture)
- #526 Allow conversation parking under --layer-split (owner-routed per-stage capture/restore)
- #525 serve: rescue a tool call stranded in an unclosed thinking span
- #523 serve: a tool call that opens inside thinking is an implicit think end
- #510 fix(serve): safely handle malformed or partial tool call arguments in history
- #504 web: saved chats in a sidebar, kept in localStorage (#361)
- #501 verify: pin the serve host thread the session loop already pins
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_S
- #488 pinned: honor STRATA_NO_LARGEPAGES on Linux; name the hugetlb pool shortfall
- #484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)
- #482 Networked Pool support
- #480 serve, web: a space-free vision temp dir, lazy vision, and a Chinese UI
- #466 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload
- #456 Monitor: the conversation cache card (parked conversations from the engine's CACHE lines)
- #454 serve: honor stop and stop_sequences
- #450 calibrate: a gpu list is not a layer split (#447)
- #441 Contrib/non mtp serving
- #440 bench: community report, RTX 5090 + Ryzen 9 9950X3D, IQ3_S at the full 262K window; prefill auto:32768 A/B; conversation-cache 32 GiB demo
- #436 serve: a non-streaming request stops when its client disconnects (#430, part of #431)
- #434 Add optional unified Strata Manager
- #426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)
- #424 fix(setup): drop an engine archive that fails to unpack
- #417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31
- #416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budget
- #409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memory
- #399 fix(setup): drop a refused engine archive instead of reusing it
- #391 iq_avx2: build the IQ2 sign table without a variable shift (BMI2 \shlx\ traps on pre-Haswell CPUs)
- #390 Multi-GPU: a --layer-split stage holds only its own layers' weights, and every card's leftover VRAM spills into a helper expert tier (4-way IQ3_S: 9,269 -> 14,172 expert slots, prefill chunk 2,048 -> 4,096)
- #385 Prompt path: a stager buffer's first job of a generation waits for the previous DMA from it (a narrow race with unpinned blobs)
- #382 HIP: MTP prompt pass per group by default (fixes prompt hang on gfx1201 with KV streaming)
- #381 Add HIP image encoder support and Linux AMD GPU telemetry
- #378 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)
- #366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Next
- #355 serve: send Cache-Control: no-cache for index.html
- #351 serve: preserve context across engine restart
- #350 Chat: a context gauge, per-request timings, and conversation compaction
- #346 setup: don't list the shared Chat settings file as a model config (KeyError: 'exe')
- #345 Add native Responses API support for Codex CLI
- #334 serve: enforce structured JSON in the native sampler
- #333 serve: add lazy startup and reliable Windows model cleanup
- #332 serve: add a standalone API request monitor
- #325 AMD working on Windows (9070XT tested)
- #321 serve: add CORS preflight (OPTIONS), reverse-proxy header support, and X-Accel-Buffering for SSE
- #314 feat: Qwen3.8-Flash-Next dense model, native Qwen3.5 attention kernel, strata-dense CLI
- #312 QSA select: ROCWMMA block-scores arm (opt-in STRATA_SELECT_WMMA=1, gfx1100)
- #309 serve: keep retained K/V through the cache capacity gate
- #302 HIP/Windows: batch expert copy dependencies in long prefill
- #288 strata-vision: --flash-attn switch, CPU runs without a CUDA context or warm-up; relative vision paths
- #284 Verify: the commit graph no longer waits; it overlaps the MTP draft
- #278 serve: count_tokens, Anthropic thinking only when asked, STRATA_REQUEST_LINES, a relative exe
- #275 Add bounded disk persistence for evicted conversations
- #269 Split prefill: run the whole prompt on the main GPU (--prefill-main)
- #257 q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact)
- #241 Faster grouped and per-hit Q2_0 expert kernels, bitwise identical
- #237 Use a second GPU's VRAM as an opt-in expert store for the prompt path
- #235 verify: STRATA_LOGPOS hook for per-position log-probabilities (items 12-13 of #84)
- #231 fix(serve): report a tool call cut off by the end of the output as unfinished
- #216 Multi-GPU: the session carve and the per-stage prompt loans
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)
- #205 fix(engine): clear stale error/cancellation string between requests and before prompt prefill
- #203 serve: prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread)
- #194 serve: clear the error string at the start of every request (fixes #183)
- #181 Windows: the engine, vision encoder and MCP servers end with the server (closing the window no longer orphans them)
- #177 Web UI: prefill speeds, request timings, a context gauge, chat compaction
- #175 Preserve independent conversation caches across interleaved requests
- #158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUs
- #154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)
- #153 Expert cache keeps ~4 GiB more on WDDM (+10% decode on a 32 GB card); Ctrl+C and client hang-ups on Windows
- #140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S support
- #139 perf(prefill): add shape-selected pre-Ampere BF16 SGEMM with current-main A/B
- #126 Windows: the expert load runs 24x slower from Task Scheduler (0.05 vs 1.42 GiB/s); a hint names the cause
- #124 Pascal (sm_60): bring the engine up on compute capability 6.0
- #122 serve: add llama.cpp-compatible discovery endpoints
- #121 Add gfx1100 HIP support with tuned prefill and bounded expert uploads
- #118 serve: routing traces from the live server
- #114 serve: report prefix reuse as usage.prompt_tokens_details.cached_tokens
- #113 serve: keep image parts in tool/assistant messages
- #111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPU
- #107 serve: report cached prompt tokens, llama.cpp timings and GET /v1/status
- #105 The Monitor's GPU readings on ROCm: an amdgpu sysfs backend, and the disk rate without psutil
- #104 The Monitor's live tok/s is a rate over the last seconds, not the mean since the first token
- #102 feat: optional RAM and VRAM guard for the server it launches
- #92 setup: name a missing or short shard with its numbers; verify reads back the manifest sha256
- #87 Turing port: run on RTX 20 (sm_75) and newer
- #84 EXPERIMENTAL Rope scaling: contexts past the trained 262K (none / linear / YaRN)
- #83 serve: llama-server-style timings on responses (prefill/decode rates for llama-swap & friends)
- #82 Web app works behind path-prefixed reverse proxies
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)
- #69 telemetry: per-request decode expert cache hit rate in DONE and web monitor
- #65 cache: preserve root system prompt checkpoint across turns and eviction
- #62 serve: the conversation cache keeps its shared prefix; the rest rotates LRU
- #61 text sampler: an empty penalty window must not touch the unsized shared bitmap
- #59 Fix/53 sampler penalties. sampler: penalties apply once, as in llama.cpp's default chain (#53); guard the unsized penalty bitmap
- #54 Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)
- #52 nvme kv cache
- #41 serve: preserve primary conversation state across auxiliary calls
- #39 telemetry: decode-only cache hit rate, prompt checkpoint retention, and dashboard
- #38 moe: continuous adaptive expert cache eviction with exponential decay
- #24 serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting
- #22 tools: add real-time web dashboard for hardware and engine monitoring
- #21 kv: add Q4_0 KV cache mode with FWHT-256 Hadamard rotation
- #19 serve: per-request temperature/top_p/top_k/min_p/penalties sampling
- #18 serve: treat unset/-1 max_tokens as remaining context, not 1024
- #16 Add optional multi-GPU expert caches and compact transfers
- #12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…
- #10 serve: read short prompt parts through the decode windows (~1 s faster first token)
- #9 CPU pool: sleep between requests instead of spinning (fixes #4)
- #8 Conversation cache: keep the chat between requests, read only what is new
- #7 serve: fix zero-token replies from engine-queue race between requests