GitHub Issues
Community · 373 Einträge
- #1373 Windows, 2-GPU layer split: 0.1.40.2 decode is 20% below 0.1.39 across the board, and STRATA_MMVQ_IL is 15% of itclosed · @Scorp1o117 · 5 Kommentare · 2026-10-08
- #1080 RTX 5000 PRO slow unsloth ud iq4_xs fixopen · @ukrolelo · 4 Kommentare · 2026-10-08
- #922 RTX 5070 Ti 16GB / Windows / IQ3_S: calibration ~14 tok/s and ~9–10 tok/s in controlled runsclosed · @annD-annD · 23 Kommentare · 2026-10-08
- #1495 [Feature]: experimental per-stage KV-grow for a two-GPU layer split at native 262Kopen · @k93k2J-glitch · 0 Kommentare · 2026-10-08
- #1468 prefill copy_i32: illegal memory access on long prompts when the prefill chunk is large (auto:32768)open · @btc1000w · 1 Kommentare · 2026-10-08
- #1444 can kvcache be saved and reused on disk?open · @badelon · 1 Kommentare · 2026-10-08
- #1061 UD-Q4_k_xlclosed · @wjxabai · 2 Kommentare · 2026-10-08
- #1488 prefill.cpp: a GCC 10 build fails (std::atomic::wait is a GCC 11 library feature, with no guard and no documented minimum)open · @zhouxihong1 · 0 Kommentare · 2026-10-08
- #1487 Windows/CUDA 11.8 Volta integration: focused patches and reproducible handoffopen · @urawazakun · 0 Kommentare · 2026-10-08
- #1485 [Bug]: update.sh cannot rebuild a locally compiled engine on a PC with a Volta card beside an RTX 40 (nvcc: Unsupported gpu architecture 'compute_70')open · @bh611 · 0 Kommentare · 2026-10-08
- #1484 [SYCL] 2x Arc Pro B70 + UD-IQ4_XS: the port's default expert kernels give corrupted output (silent wrong answers); STRATA_EXPERT_SPLIT=1 restores itopen · @friedrichAl · 0 Kommentare · 2026-10-08
- #1483 Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm buildingopen · @1314521gjy · 2 Kommentare · 2026-10-08
- #1431 Feature request: tool-call emission recovery for low-bit quants in agent workloads (Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS)open · @mechanicss · 1 Kommentare · 2026-10-08
- #1482 [Feature]: publish a ready-made Docker image, so users do not compile the engine themselvesopen · @nullata · 0 Kommentare · 2026-10-08
- #1479 HIP: STRATA_HEAD_MIX_MULTI / STRATA_ONE_TOKEN_COMMIT are hard-off; as opt-ins head-mix is exact and +0.4% decode on R9700 (patch inside)open · @jkuepker · 0 Kommentare · 2026-10-08
- #1478 gfx1201 (R9700): the gfx1151 prompt switches in arch_defaults.cpp are exact and +4-8% prompt; suggest a gfx12 entry (decode half is slower)open · @jkuepker · 0 Kommentare · 2026-10-08
- #1409 What is normal prefill tok/s? Suggest to publish this infoclosed · @markd89 · 0 Kommentare · 2026-10-08
- #921 CPU expert-pool 20 ms spin can hurt GPU-heavy decode; 100 µs gives ~15% higher TG and ~4× lower CPU usage on dual RTX 4090open · @hipotures · 4 Kommentare · 2026-10-08
- #1397 SYCL port on Windows/OpenCL: 0.1.40-sycl decodes at 22.3 tok/s with --spec 2 where 0.1.39-sycl does 53.6 (and short runs vary 40%)open · @demetree · 2 Kommentare · 2026-10-08
- #1386 2x RX 7900 XTX (gfx1100, layer split): STRATA_PF_GEMM is worth ~7% prefill and is off by defaultopen · @yalu-arch · 1 Kommentare · 2026-10-08
- #1473 STRATA_SYCL_SPIN_MAX default: a JIT build of the B70 gets the A-series bound and loses 46-67% of decodeopen · @demetree · 1 Kommentare · 2026-10-08
- #872 Session RESTARTS COMPUTER after 25-30 mins (gfx1100)open · @vletu-ghub · 9 Kommentare · 2026-10-08
- #1475 expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)open · @jkuepker · 0 Kommentare · 2026-10-08
- #1474 hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7open · @jkuepker · 0 Kommentare · 2026-10-08
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAMopen · @lollo78 · 9 Kommentare · 2026-10-08
- #1469 0.1.40.2: decode on a Tesla P40 (sm_61) is about half of 0.1.40's; bisected to 2e4ddf6, and restoring __restrict__ in pdl.hpp (STRATA_PDL_RESTRICT) gives back most of itopen · @paulhothersall · 2 Kommentare · 2026-10-08
- #1445 vision encoder did not start on V100-SXM2-32GBclosed · @TAIIOK · 2 Kommentare · 2026-10-08
- #1352 Multi-GPU performance is ~2x lower than Francesco Albano fork despite the changes being merged into v0.1.40open · @chemical12 · 5 Kommentare · 2026-10-08
- #1440 [Bug]: [SYCL] 0.1.40.2 on 2x Arc Pro B70: --layer-split fails three ways (#1054's fix from #1111 is not in main, plus two new split regressions)open · @crobe201 · 2 Kommentare · 2026-10-08
- #1422 Modelo Qwen3.8-flash-next? Alguém mais percebeu que parecem mais com Qwen3.6!open · @scsrat · 5 Kommentare · 2026-10-07
- #1118 Batch slots: the engine stalls ("no progress for 60 s … reading the prompt (batched)") when a long prompt is read while another slot decodes; STRATA_BATCH_DECODE_SHARE=0 avoids itclosed · @uncle-daddy-jp · 3 Kommentare · 2026-10-07
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA archclosed · @lineape · 1 Kommentare · 2026-10-07
- #1404 [Feature]: run the image encoder on a second PCopen · @TraceRecursion · 2 Kommentare · 2026-10-07
- #1466 60G VRAM (3090+3060*3), 256K context, 75~85 T/Sopen · @ee-derrick · 1 Kommentare · 2026-10-07
- #847 Allow system ram to be used instead of SSD when available.closed · @jarrodhroberson · 12 Kommentare · 2026-10-07
- #1356 Stream corruption with `STRATA_QFUSE=1` on GFX1151 when using `parallel: 2`open · @ChrisDeadman · 2 Kommentare · 2026-10-07
- #612 Possible to support gfx1151?open · @wolf0403 · 8 Kommentare · 2026-10-07
- #1464 Add vision to docker containeropen · @ceinstaller · 0 Kommentare · 2026-10-07
- #1463 MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)open · @mathcuei · 0 Kommentare · 2026-10-07
- #723 Design question before we build: two lanes on a peer pair that help each other ("mutual help") — welcome upstream, or keep it in a fork?open · @q8atnight · 2 Kommentare · 2026-10-07
- #1453 Live agent workload on RX 7900 XTX (24 GB) + 30 GB RAM — 70–74 tok/s, 91–96% expert cache hitsopen · @artfix · 0 Kommentare · 2026-10-07
- #1387 [Feature Request] Support for higher quantizations (Q8 / unquantized) for Qwen 3.8 / 4 with high-end hardware (128GB VRAM / 192GB RAM)open · @Bobsosa · 2 Kommentare · 2026-10-07
- #1423 Support for AMD Radeon 780M, Phoenix , gfx1103open · @Paco1960 · 1 Kommentare · 2026-10-07
- #1447 docs: MULTI_GPU.md says --pipeline-windows 2 and --adapt-async 1 exclude each other, but they combineopen · @adambenhassen · 0 Kommentare · 2026-10-07
- #1121 --batch-mtp kills the engine on an IQ3_S (GSQ-RCO) pack: mtp: unsupported native MMVQ GGML typeopen · @Benderyu · 7 Kommentare · 2026-10-07
- #558 Coder is no good, worse than 3.8 27B 4 bitclosed · @frederikhors · 5 Kommentare · 2026-10-07
- #1442 Feature request: support chat title generating from Xcode's agentopen · @hys17 · 0 Kommentare · 2026-10-07
- #1054 [SYCL / Intel Arc] 2x Arc Pro B70 --layer-split auto deadlocks at startup after the host-mirror fillopen · @djbrettb · 3 Kommentare · 2026-10-07
- #937 `[gfx1030] HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION in gdn_step_commit_kernel after a few long fresh prefills`open · @lucaravera-glitch · 2 Kommentare · 2026-10-07
- #1180 STRATA_PF_FUSED on gfx1100: missing __syncthreads() in native_w11_kernel + waves_per_eu(8) wrong on clang 22 / ROCm 7.10open · @zhhhn · 2 Kommentare · 2026-10-07
- #1267 gfx1151 (Strix Halo) gets the ROCm 7.10.0a20251120 wheels, not the 7.14.1 that docs/STRIX_HALO.md was measured on: the engine crashes on kernel 7.2.8, and prompts run at half speedclosed · @alexdns1 · 4 Kommentare · 2026-10-07
- #1412 Windows: `SeLockMemoryPrivilege` moves the expert arena to 2 MiB large pages — but the token is issued at logon (1314 ≠ 1450)open · @shahrokhzargarpour · 1 Kommentare · 2026-10-07
- #1428 --layer-split auto on a P40 + RTX 3070 (3070 first) flips from K=43 to K=47 with --kv int8 or --vram-reserve-mib 600, and prefill drops from 458 to 111-113 tok/s (0.1.40.2)open · @paulhothersall · 0 Kommentare · 2026-10-07
- #929 `--batch` with `--layer-split` exits ("verify batch: layer 34 never rang") with a 3070 first and a P40 second; works with the P40 first (CUDA, v0.1.39)open · @paulhothersall · 3 Kommentare · 2026-10-07
- #1238 0.1.40: --layer-split auto with STRATA_STAGE_TRIM=1 still picks K=2 on a P40 + RTX 3070; 572b607 fixes it but is not in the releaseopen · @paulhothersall · 2 Kommentare · 2026-10-07
- #1146 LLM Server feature request: "idle_load: true" parameteropen · @KoenKeunen · 2 Kommentare · 2026-10-07
- #1425 [Windows, RTX 5090 32 GB, NVFP4 fork] ple: --ple-io direct streams the n-gram table at ~34 MB/s while the SSD does 1.4 GB/s - fresh long prompts prefill ~5x slow; cold passes can trip the #29 watchdogopen · @RichardFreml · 0 Kommentare · 2026-10-07
- #1005 CRLF line-break tokens (`\r\n`, ids 317/845/8301) get far too little probability; everything else matches llama.cppopen · @brenoperucchi · 3 Kommentare · 2026-10-07
- #1233 ``` Windows + Strix Halo (gfx1151, Radeon 8060S): UD-IQ4_XS runs end-to-end — decode flat 29.5-30.4 tok/s at 1K-32K, but a GPU driver timeout (TDR) during a cold 32K prefill ```open · @Facostarr · 4 Kommentare · 2026-10-07
- #782 Codex got error "input items of type 'additional_tools' are not supported"closed · @Anivie · 4 Kommentare · 2026-10-07
- #1203 AMD Radeon Pro V620 (gfx1030) - Dockerfile ???open · @luckylinux · 2 Kommentare · 2026-10-07
- #1392 Web app: an assistant reply with no content is dropped from the conversation history (reasoning models repeat themselves)closed · @demetree · 1 Kommentare · 2026-10-07
- #1385 Intermittent 'int' object has no attribute 'ranks' in tokenizeropen · @mustafaerdinc · 1 Kommentare · 2026-10-07
- #1380 No data shown for 9070xt on dashboard.open · @Roshanlikeocean · 1 Kommentare · 2026-10-07
- #1357 mtp: the per-token drafter router call is missing the n_expert == 512 guard its three sibling call sites carryclosed · @rwkeyes · 1 Kommentare · 2026-10-07
- #1244 Docker v0.1.40: persisted /data/config edits ignored on restart; stale /opt/strata config used (worked in v0.1.39)closed · @ampersandru · 3 Kommentare · 2026-10-07
- #1194 `file_cache_keeps()` picks unbuffered (O_DIRECT) expert reads on a 32 GB + 12 GB GPU box with `--resident-budget-gib`, halving decode; `STRATA_UNBUFFERED_LOAD=0` restores itclosed · @willthehuman · 4 Kommentare · 2026-10-07
- #1070 New Feature - add option to change location of modelclosed · @Gifciak · 5 Kommentare · 2026-10-07
- #1413 `--batch-groups G>1` + `--batch-mtp` is silently inert: the pipelined path never calls `batch_step`, but still builds per-slot draftersopen · @shahrokhzargarpour · 0 Kommentare · 2026-10-07
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTEDopen · @galmok · 7 Kommentare · 2026-10-07
- #1139 GDN batch path skips activation quantization when STRATA_QFUSE=1closed · @pepuscz · 5 Kommentare · 2026-10-07
- #1410 Support Intel Arc Pro B50 (Xe2 / BMG-G21) as a tested SYCL target — 16 GB at 70 W is the card users are actually buyingopen · @ditronicos · 0 Kommentare · 2026-10-07
- #1407 [gfx1200, 2x RX 9060 XT layer split] prompt-read stalls (watchdog #29 aborts) + run-to-run prefill degradation: fresh prompts slow 3-4x mid-run, file tier re-reads 24-35 GB per requestopen · @Iokkeh2025 · 0 Kommentare · 2026-10-07
- #1403 Almost no changes from 2 recent updates,open · @Noctorngaming · 0 Kommentare · 2026-10-07
- #1206 DirectStorage implementationopen · @notGabeNewell · 2 Kommentare · 2026-10-07
- #771 Linux expert arena (0.1.39): MADV_HUGEPAGE + defrag=madvise makes the arena's faults 20x slower (25 s start -> 432 s), and 'loaded ... at X GiB/s' does not show itclosed · @Alucard24 · 6 Kommentare · 2026-10-07
- #692 Feature Request: Intel Arc A770 Supportopen · @wyt2439843587-crypto · 8 Kommentare · 2026-10-07
- #1393 Support for AMD Radeon 890M APUopen · @rutravis · 0 Kommentare · 2026-10-07
- #857 "parallel": 2 lowers total throughput on a single 48 GB card (100% expert-cache hits): keep MTP drafts in batch windows?closed · @bobvious · 5 Kommentare · 2026-10-07
- #1389 RX 7900 XTX (gfx1100): 925 -> 2,503 tok/s prompt read (+171%) - the method, and the 26% cliff below full lend coverageopen · @zhhhn · 0 Kommentare · 2026-10-07
- #906 Older PC (Haswell, DDR3, PCIe 3.0 x8): +15% decode on Q2_0 and IQ3_S from #706, #764 and a busier adaptive tieropen · @1872183316 · 1 Kommentare · 2026-10-07
- #1348 Foresight: fetching MoE experts before they're needed — measured on 2× 3090 (update: prefetch path built and correct, no decode gain yet; data, tools, help wanted)open · @q8atnight · 1 Kommentare · 2026-10-07
- #713 Proposal: a common benchmark format and a comparison page for community reportsopen · @brenoperucchi · 4 Kommentare · 2026-10-07
- #1276 UPDATE.bat cannot update existing checkouts after the main history rewriteclosed · @nirvash · 3 Kommentare · 2026-10-07
- #1341 serve: a >100k-token prompt deadlocks the pipelined verify-window reader (--pipeline-windows >= 1) - the #29 watchdog aborts the engine and the server reloads the modelopen · @btc1000w · 2 Kommentare · 2026-10-07
- #657 Error while starting - PLE table size mismatchclosed · @alexit190 · 4 Kommentare · 2026-10-07
- #1128 0.1.40 --kv-grow: a short request trims the K/V to 8192 cells and drops the conversation cache, so the next 128K+ continuation is re-read in full (20.9% of continuations, RTX 5090, IQ3_S)open · @Darkstarrd-dev · 2 Kommentare · 2026-10-07
- #968 0.1.39: UD-IQ4_XS prompts read as garbage on an RTX 5090 (sm_120) through the MMQ prompt path; STRATA_PREFILL_MMQ=0 or an sm_86 card reads them correctlyopen · @signalnine · 6 Kommentare · 2026-10-07
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)closed · @maverde73 · 3 Kommentare · 2026-10-07
- #1375 Strata on an RX 6600 XT (gfx1032, 8 GB) and a Xeon E5-2678 v3 with DDR3open · @Stataltex · 0 Kommentare · 2026-10-07
- #1371 [ok] strata-swift-iq3_xxs x3open · @XeonG · 1 Kommentare · 2026-10-07
- #875 P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)open · @paulhothersall · 4 Kommentare · 2026-10-07
- #1214 Vulkan implementation for multi-brands setupsclosed · @notGabeNewell · 2 Kommentare · 2026-10-07
- #1369 Two 32K prompts on one RTX 5090: overlapping them matched back-to-back wall timeopen · @steve8697 · 0 Kommentare · 2026-10-07
- #1366 [AMD/gfx1031/Windows] Self-built engine crashes with an access violation inside amdhip64_7.dll before any outputopen · @zunebakananu71-alt · 0 Kommentare · 2026-10-07
- #967 Unsloth ud-q4_k_xl - no vision?closed · @kelheor · 9 Kommentare · 2026-10-07
- #951 crash after first inputsopen · @Maelstrom2014 · 2 Kommentare · 2026-10-07
- #672 Better multiGPU supportclosed · @sudoeste · 2 Kommentare · 2026-10-07
- #1364 Metrics logs & live dashboardopen · @skyprince999 · 0 Kommentare · 2026-10-07
- #932 Dual GPU supportopen · @runningman84 · 9 Kommentare · 2026-10-07
- #1280 bench: community report, 2x RTX 3090 (Linux, Docker): 0.1.40 resident RAM mode on a split (#848) and --pipeline-windows (#859)open · @andrej-reimer · 1 Kommentare · 2026-10-07
- #1363 3x 5060ti 16gb crashopen · @fabionguyen91 · 0 Kommentare · 2026-10-07
- #1317 serve: an untimed read of the engine/vision pipe can hold the request turn forever (intermittent hang; the #481 lost-step class)open · @btc1000w · 1 Kommentare · 2026-10-07
- #1304 Prefill/Reading the prompt with multi-gpu cards is slowopen · @sense1024 · 2 Kommentare · 2026-10-07
- #1302 0.1.40 SYCL port on 2x Arc Pro B65: expert-cache fill stalls (1-core spin + xe "Timedout job" + device coredump); 0.1.38-based engine works on the same cardopen · @james151br · 1 Kommentare · 2026-10-07
- #1275 Windows HIP: --prefill auto can pick a chunk whose VRAM borrow leaves too little for the verify captures — silent engine exit at serve initopen · @slogomansdad · 1 Kommentare · 2026-10-07
- #1272 0.1.40: with the shared-expert stream fork off by default, decode on 2x RX 6900 XT (gfx1030, Linux) is 5-13% slower; STRATA_SH_STREAM=1 restores it, and the #884 stall does not come back with itopen · @xjc10 · 1 Kommentare · 2026-10-07
- #1261 461open · @maraa081 · 1 Kommentare · 2026-10-07
- #1256 Feature Request: Support for RPCsclosed · @cumal · 1 Kommentare · 2026-10-07
- #1252 DraftPolicy: a lookup window size priced during a slow stretch stays over-priced for the rest of the processclosed · @guthirry · 1 Kommentare · 2026-10-07
- #1250 [Linux][RTX 3080 12 GB, 62 GB RAM] 0.1.40.1 start OOMs the host while allocating the page-locked expert copy (IQ2_XS); 0.1.31 starts fineopen · @Absolutely-not-a-coder · 6 Kommentare · 2026-10-07
- #1229 serve: phone photos reach the model rotated by 90° (EXIF Orientation is ignored)closed · @crazyaimachine · 1 Kommentare · 2026-10-07
- #1217 gfx1150 (Radeon 890M / Strix Point) works: report and a one-line CMake patchclosed · @zeriyoshi · 1 Kommentare · 2026-10-07
- #1191 Automatic support of other variants and quantsclosed · @kelheor · 1 Kommentare · 2026-10-07
- #1188 --kv k8v4 produces degenerate repetition when --kv-resident streams (0.1.40)closed · @baiqvesse · 3 Kommentare · 2026-10-07
- #1169 k8v4 + --kv-resident: the batched KV append (#783) skips the host mirror for hybrid — parking and KV streaming serve stale bytes (cross-conversation contamination on 0.1.40)closed · @taxah92 · 1 Kommentare · 2026-10-07
- #1135 k8v4: deterministic token repetition on sequential generation (RTX 5090, 0.1.40) — byte-identical across two quantsclosed · @gravitomagnetic · 4 Kommentare · 2026-10-07
- #1116 major performance drop in latest 0.40 engine -kinda fixedopen · @XeonG · 5 Kommentare · 2026-10-07
- #1103 hip: gfx1030 (RX 6800 XT) intermittent verify timeouts (#267): ROCm 7.14 gfx103X wheels regression, fixed by 7.13open · @ChenCui26 · 2 Kommentare · 2026-10-07
- #1094 --layer-split auto fails on dual GPUs for 512k contextopen · @fukc-gihtub · 1 Kommentare · 2026-10-07
- #1079 bench: community report, V100 32 GB + P100 16 GB (sm_70 + sm_60, Linux source build): IQ3_XXS at 262K with the P100 as a helper expert cache, 0.1.39 vs 0.1.40closed · @aleesposito85 · 2 Kommentare · 2026-10-07
- #890 HIP: "parallel": 2 + --layer-split on 2x R9700 exits (code 1) after "captured the batch window over slots 0,1"open · @asaffulks · 6 Kommentare · 2026-10-07
- #1074 Build fails on openEuler 24.03 / GCC 12.3 (sm_70): vmm.hpp uses an unqualified size_tclosed · @BaiQichen2026 · 2 Kommentare · 2026-10-07
- #1073 Wrench Thailand and Drill Bits Thailand: Reliable Hand and Power Tool Solutions from M10 TOOLSclosed · @m10hardware25-code · 1 Kommentare · 2026-10-07
- #1072 Vision: >64 unique images can evict .sve files still referenced by the current requestclosed · @jinyu0 · 2 Kommentare · 2026-10-07
- #1071 vmm.cpp: the CUDA 12 engine does not build with a toolkit older than 12.5 (missing CUDART_VERSION guard)closed · @zhouxihong1 · 1 Kommentare · 2026-10-07
- #1069 Tesla P100 (Pascal, sm_60) + RTX 2070 SUPER, Linux/Docker (CUDA 12.9), 32 GB RAM: Q2_0 low-RAM mode, about 30 tok/s decode on the P100 alone and 36-38 tok/s with the second card as a helper expert cache (benchmark.py)closed · @tobby1968 · 1 Kommentare · 2026-10-07
- #1065 9070xt gpu not detectedopen · @chicco83 · 1 Kommentare · 2026-10-07
- #1059 serve: a request the engine rejects with ERR waits the full 300 s before failing (STOP + _drain_control("DONE") after an ERR), holding the control lockclosed · @cardosofelipe · 1 Kommentare · 2026-10-07
- #1056 Prefill: Stager threads yield-spin while waiting. On Linux with --mmap-experts (DGX Spark / GB10), prompts are read ~10x slower; elsewhere ~20 cores are burned during prefillclosed · @uncle-daddy-jp · 1 Kommentare · 2026-10-07
- #1053 A reply that ends inside `<think>` (no `</think>`) comes back empty with `finish_reason: "stop"`, and the agent stops — close the thinking and continue onceclosed · @davidsheridan77-dot · 3 Kommentare · 2026-10-07
- #1052 网赌被黑后如何正确应对?咨询Q398722520closed · @djeiywuwy7627267262-cyber · 4 Kommentare · 2026-10-07
- #1051 Website Development Company in Jaipur: Why Sangita Technologies Is the Smart Choice for Your Businesclosed · @sangitatechnologies · 1 Kommentare · 2026-10-07
- #1034 vram_elastic / POST /v1/vram: response schema and failure semantics (Windows 11, RTX 4090)open · @oliver-prog-0705 · 3 Kommentare · 2026-10-07
- #1029 docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)closed · @sarge18 · 1 Kommentare · 2026-10-07
- #1026 zakoclosed · @zakoz123191-del · 1 Kommentare · 2026-10-07
- #1015 Reasoning stuck in loop on Coder modelopen · @brishisharma · 1 Kommentare · 2026-10-07
- #1008 Title: 0.1.39 decode is ~15–23% slower than 0.1.38 on AMD (gfx1201/HIP, 2× R9700); prefill unchangedclosed · @uptou888 · 3 Kommentare · 2026-10-07
- #997 0.1.39: engine exits (code 1) with "parallel": 4 on UD-IQ4_XS, RTX 3090 24 GB, three concurrent 8K prompts (reproducible; IQ3_S unaffected)open · @cardosofelipe · 5 Kommentare · 2026-10-07
- #990 MCP install planning rejects supported AMD CPU vision (and retains the old Windows HIP restriction)closed · @kapelame · 1 Kommentare · 2026-10-07
- #989 Integrating for Open Source LLM Hardware Support with GSys LibreCoreclosed · @etcimon · 1 Kommentare · 2026-10-07
- #986 Zed says Tools Unsupported: `/props` has no `chat_template_caps`closed · @1it · 1 Kommentare · 2026-10-07
- #985 Also allow Visual Studio 2026 to build from sourceclosed · @vehystrix · 1 Kommentare · 2026-10-07
- #984 reasoning_budget_tokens is ignored?open · @PentaForgeDev · 5 Kommentare · 2026-10-07
- #980 Field report: local Qwen3.8-Flash-Next (Strata) as an agent building 3D scenes in Unreal Engine 5.8 via MCP — what broke, what fixed it, what is still missingopen · @talisp · 1 Kommentare · 2026-10-07
- #978 Strata 7900xtx Qwen 3.8 Flash Next Q3_xxs VS Openai gpt6 Lunaclosed · @2jztricks · 1 Kommentare · 2026-10-07
- #964 SM75 (2080Ti) dual-GPU: "layer X never rang (illegal memory access)" and verify-window hang on IQ3_Sopen · @dehuiyede · 2 Kommentare · 2026-10-07
- #961 Windows (WDDM): verify window wedges nvlddmkm — 0x141 engine timeout → 0x116 TDR failure (driver defect, not a Strata bug)closed · @1593914054 · 1 Kommentare · 2026-10-07
- #959 ./setup.sh fails on Gentoo Linuxopen · @kyeno · 5 Kommentare · 2026-10-07
- #957 using omp, pi with strata bad results.closed · @Almazick · 3 Kommentare · 2026-10-07
- #954 [BUG] 0.1.39: fused int8 prompt experts (STRATA_PF_FUSED=1) + #583 byte-budget ring → kernel hang, `prefill mmq: iota: unknown error` on sm_120 (RTX 5090, WSL2)open · @Krypto-Whitehat · 3 Kommentare · 2026-10-07
- #953 Windows AMD report: UD-IQ4_XS on RX 7900 XTX (gfx1100) works, 34-64 tok/s decodeopen · @Huobi-cloud · 2 Kommentare · 2026-10-07
- #948 A40 - failed to runopen · @catamaican · 1 Kommentare · 2026-10-07
- #942 AMD `gfx1102 / RX 7600 XT` - successful real model run on Strata 0.1.39open · @Lefox-DeMod · 2 Kommentare · 2026-10-07
- #938 rx 7600xt 16gb supportclosed · @TR3YV3N · 1 Kommentare · 2026-10-07
- #935 The models are subpar.closed · @ErfolgreichCharismatisch · 28 Kommentare · 2026-10-07
- #926 一直Thinking,不干事closed · @CNWenwuGong · 5 Kommentare · 2026-10-07
- #918 Windows: Radeon 8060S (gfx1151) builds, selftests and passes the HIP ctest on top of #895open · @storm-ace · 1 Kommentare · 2026-10-07
- #915 Windows AMD report: RX 6800M (gfx1031, laptop) works with a self-built HIP engine — please add 0x73DF / gfx1031 to the Windows pathopen · @Ultmate-hub · 1 Kommentare · 2026-10-07
- #909 Proposal (docs only): a short "contributing a change or report" section for AGENTS.md, and a pinned list of test requests. Yes/no?open · @paulhothersall · 1 Kommentare · 2026-10-07
- #899 Are there plans to support Apple Silicon as well?closed · @pangty06116 · 1 Kommentare · 2026-10-07
- #893 serve/server.py reads POST bodies as empty when sent with Transfer-Encoding: chunked (no Content-Length fallback) — breaks requests through some reverse-proxy/relay agentsclosed · @oscar-investmatic · 1 Kommentare · 2026-10-07
- #892 Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 worksopen · @Thanh-Huy1104 · 2 Kommentare · 2026-10-07
- #891 SYCL port v0.1.39 fails to compile: session.cpp ThreadAffinity type errorsclosed · @asaffulks · 2 Kommentare · 2026-10-07
- #884 HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)open · @xjc10 · 6 Kommentare · 2026-10-07
- #879 All-logits-NaN degeneration (one repeated token forever) still fires on 0.1.39b / current main — first poison traced to the MoE output rows: garbage bits, varying layeropen · @66419118nnn · 5 Kommentare · 2026-10-07
- #870 Intel Arc Pro B60 (e211), 24 GB: first run on 0.1.39 + 4 fixes for current mainclosed · @Magh97 · 2 Kommentare · 2026-10-07
- #867 Intel SYCL port on 2x Arc Pro B60: results, and "never rang (graph finished)" on the host-mirror pathclosed · @LocalXPU · 4 Kommentare · 2026-10-07
- #862 Got it working on my old 4 Maxwell Titans X, and 2016 dated Supermicro 4029TRT for Qwen 3.8 Flash Next Q8_0closed · @phalexo · 3 Kommentare · 2026-10-07
- #842 cache files problem and other one problemclosed · @gitforestit · 1 Kommentare · 2026-10-07
- #841 not stable start with 3+ cardsopen · @yotadrivers-blip · 8 Kommentare · 2026-10-07
- #836 Milestone: Add Apple Silicon Cross-Platform Supportclosed · @Atlessc · 1 Kommentare · 2026-10-07
- #831 The profile's length caps the expert arena — even an explicit --expert-cache Nclosed · @ZhongUncle · 2 Kommentare · 2026-10-07
- #828 Heap corruption (munmap_chunk: invalid pointer) at startup in Prefill::init — crashes with 1/2/4 GPUsopen · @wehooper4 · 3 Kommentare · 2026-10-07
- #825 Results for --model UD-IQ4_XSclosed · @Almazick · 4 Kommentare · 2026-10-07
- #817 Assistant turn can be interrupted when generation emits a reserved ChatML boundary markeropen · @d4mer · 2 Kommentare · 2026-10-07
- #814 Would it be possible to add support for the AMD Radeon RX 6750 GRE 12GB?closed · @lzb379440026 · 1 Kommentare · 2026-10-07
- #806 Strange behavior for x2 RTX 5060 Ti (16Gb each) and UD-Q4_K_XLclosed · @art-den · 1 Kommentare · 2026-10-07
- #791 Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer splitclosed · @yy16432 · 2 Kommentare · 2026-10-07
- #781 1M context: identical expert work per window, but the VRAM-touching stages grow 7-8xopen · @1314521gjy · 7 Kommentare · 2026-10-07
- #779 laptop crash with latestopen · @XeonG · 3 Kommentare · 2026-10-07
- #774 Support for Nvidia GB10.open · @geniousdahiya · 2 Kommentare · 2026-10-07
- #772 Community project: strata-router, a multi-node router for Strata / 社区项目:Strata 多节点路由入口closed · @konijiwa110 · 1 Kommentare · 2026-10-07
- #770 Qwen 3.8 27B on Strata?closed · @frederikhors · 3 Kommentare · 2026-10-07
- #765 Make VRAM budgeting aware of KV, prefill staging and expert residencyopen · @j-luwierski · 4 Kommentare · 2026-10-07
- #760 524K via rope-scaling on consumer Blackwell (sm_120, no clusters): the histogram fallback works end to end — field data + two budget gotchas on 16GB cardsclosed · @johnmosesventura06 · 3 Kommentare · 2026-10-07
- #754 Anthropic API: tool calls intermittently emitted inside `thinking` and returned as `end_turn`closed · @JakeCow · 1 Kommentare · 2026-10-07
- #746 Regarding 128GB Memory Optimization Strategiesclosed · @Susan985 · 5 Kommentare · 2026-10-07
- #739 [enhancement] Variance-normalized KV-cache quantization (KVarN) & KV cache precision tailopen · @gregfromo · 1 Kommentare · 2026-10-07
- #736 5090 Laptop 24G + 96 GB RAM swift-iq3_xxs和qwen-iq3_s 524K 是最优档:decode 完全不掉closed · @akiry09 · 1 Kommentare · 2026-10-07
- #735 can't load vision encoderclosed · @lanny0914-2 · 2 Kommentare · 2026-10-07
- #729 KV int8 vs fp16 KLD compare at 240K contextclosed · @jerry78424 · 2 Kommentare · 2026-10-07
- #728 Infinite loop issueopen · @davidliudev · 18 Kommentare · 2026-10-07
- #714 5090 Laptop 24G + 96 GB RAM Unsloth UD-Q4_K_XL测试结果closed · @akiry09 · 6 Kommentare · 2026-10-07
- #710 v0.1.38: cross-request tool loop repeats an ineffective repair 98 times despite a thinking budgetclosed · @Traveller23 · 5 Kommentare · 2026-10-07
- #709 Does the HIP backend work on non-AMD GPUs via chipStar (HIP → SPIR-V)? Spike says yes with a 3-line shim diffopen · @mq-beefcake · 1 Kommentare · 2026-10-07
- #703 [AMD HIP] RX 6750 XT (gfx1031) works in heterogeneous layer split; long multimodal prompts can crash the engineclosed · @PJGV333 · 1 Kommentare · 2026-10-07
- #702 [AMD HIP] RX 9060 XT (gfx1200) Repeating outputclosed · @Gerporgl · 5 Kommentare · 2026-10-07
- #694 fit_max_tokens issue, context problem with agentic coding.closed · @mkultra333 · 3 Kommentare · 2026-10-07
- #679 Feature Request: Load/Offload/Swap mmproj dynamically from RAM to VRAM when usedclosed · @Interpause · 12 Kommentare · 2026-10-07
- #678 Feature request: Optional Intel NPU support for the vision encoderclosed · @eveloso-web · 1 Kommentare · 2026-10-07
- #669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scoresopen · @gputier · 4 Kommentare · 2026-10-07
- #658 Proposal: single-GPU prefill buffer planning and RAM budgeting for conversation caching (RTX 5090 measurements)open · @spideytznn · 1 Kommentare · 2026-10-07
- #654 HIP (gfx1200) RuntimeError Exception Code: 0xC0000005open · @remielowik · 10 Kommentare · 2026-10-07
- #642 A dual-GPU fork for 32 GB PCs (resident mode on a layer split, pipelined windows)closed · @Hardin22 · 10 Kommentare · 2026-10-07
- #625 Vision (CPU path): offer a Q8_0 mmproj and a higher image-token capclosed · @Thxeverybody · 4 Kommentare · 2026-10-07
- #606 After a 36,689-token request at 155k, every later request answers one repeated token until the engine reloadsopen · @66419118nnn · 7 Kommentare · 2026-10-07
- #579 AMD/HIP RX 9070 XT (gfx1201) stalls during batched prompt ingestion on ROCm 10.0open · @s91402001 · 7 Kommentare · 2026-10-07
- #541 AMD/HIP (R9700, engine 0.1.36): batched prompt read stalls at a repeatable token; STRATA_PLE_BATCH=0 avoids itclosed · @icodebot · 2 Kommentare · 2026-10-07
- #535 Thinking speed suddenly drops to 0.5 token/sclosed · @sinand99 · 2 Kommentare · 2026-10-07
- #528 Conversation cache: decode drops ~4-6x on continued conversations (0.1.36, Windows, RTX 5090, IQ3_XXS)closed · @NikolaFC · 4 Kommentare · 2026-10-07
- #519 0.1.36 on an RTX 5090: STRATA_PF_FUSED=1 speeds up IQ3_XXS prompts 13-23%, and a ~550-token prompt spends ~470 ms in the batched prompt pathopen · @brenoperucchi · 7 Kommentare · 2026-10-07
- #517 new error `raise RuntimeError("the engine exited before it was ready" + (f" (see {log})" if log else "") +`closed · @XeonG · 6 Kommentare · 2026-10-07
- #1353 stager-transient-regressionopen · @dag08 · 0 Kommentare · 2026-10-07
- #1350 Wouldn't it be better if the server is C++ too?open · @TByte007 · 0 Kommentare · 2026-10-07
- #1349 2× RTX 2080 Ti (Turing, NVLink) field report: DDR4-2666→3600 memory A/B, own vs borrow vs peer-device by context size, three-switch stackopen · @zyYuc · 0 Kommentare · 2026-10-07
- #1347 serve: a full conversation cache drops the snapshot at the physical-RAM gate instead of evicting to pass it (270k tokens re-read, 96 s, 32 times)open · @alanthinker · 0 Kommentare · 2026-10-07
- #1346 __permission_probe__closed · @alanthinker · 0 Kommentare · 2026-10-07
- #1344 Server: Multiple API key(llamacpp format) will seen as one long keyopen · @ed-hch · 0 Kommentare · 2026-10-07
- #1337 setup --calibrate loses the measurements it already has when a later engine start failsopen · @aly8246 · 0 Kommentare · 2026-10-07
- #1285 Dual RTX 5090: Strata 0.1.40 production decode at 243 tok/s and synthetic peer-prefill resultsclosed · @b1naryagent · 0 Kommentare · 2026-10-07
- #1258 gfx1151 (Strix Halo) prompt: per-kernel profile of a 16K prompt (rocprofv3) - where would outside help be useful?open · @huppiflupp · 2 Kommentare · 2026-10-07
- #1322 setup re-run with --vision gpu does not add --vision to an existing config's argsopen · @arthurlin1979 · 0 Kommentare · 2026-10-07
- #1312 [Windows][RTX 3090 24 GB, 64 GB RAM, IQ3_S] Feature request: expose the QSA indexer budget (512 blocks / 2048 tokens)open · @eakkawat · 0 Kommentare · 2026-10-07
- #1248 SYCL build fails on Intel Arc A770 (Alchemist) - v0.1.40.1open · @Javalopes · 1 Kommentare · 2026-10-07
- #1308 Traitement de l'eau et rondelle de roue : les solutions industrielles de Tecnoter Groupopen · @tecnotergroup2-a11y · 0 Kommentare · 2026-10-07
- #1306 Strata currently has incomplete support for ARM64 (particularly GB10).open · @jeasinlee · 0 Kommentare · 2026-10-07
- #1301 docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating textopen · @ZackO2o · 0 Kommentare · 2026-10-07
- #1300 bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes fromopen · @ZackO2o · 0 Kommentare · 2026-10-07
- #1299 AMD vison encoder on Windowsopen · @Zethalion · 0 Kommentare · 2026-10-07
- #1145 Resident RAM mode on a layer split — 3 GPUs, UD-IQ4_XS, 40 GB RAM (engine 0.1.40)open · @chiangww · 2 Kommentare · 2026-10-07
- #566 Proposal: extend existing calibration to Linux HIPopen · @xyzzing · 3 Kommentare · 2026-10-07
- #1294 有没有docker部署的镜像包open · @wbdy1 · 0 Kommentare · 2026-10-07
- #1286 PCIe Issue with Strata and RTX 3060open · @clawdsrlb-hub · 0 Kommentare · 2026-10-07
- #1129 New Feature - Add option so pick the default model settings: --thinking --instructopen · @victorgabr · 1 Kommentare · 2026-10-06
- #1277 HIP: fused IQ expert kernels via RDNA4 int8 WMMA (RX 9060 XT / gfx1200)open · @slogomansdad · 0 Kommentare · 2026-10-06
- #816 0.1.39: decode on an RX 6800 (gfx1030, Windows) is 12-22% slower than 0.1.38; STRATA_SH_STREAM=0 removes itclosed · @did-technomancer · 12 Kommentare · 2026-10-06
- #690 Layer split: native prefill kernels fail on the second GPU (no kernel image available), regardless of which card it isopen · @Nauclerus · 2 Kommentare · 2026-10-06
- #1178 serve: Special tokens (<|im_end|>) emitted during model thinking cause tasks to terminate earlyopen · @TwoAnts · 5 Kommentare · 2026-10-06
- #1257 Feature request: A separate "copy" button at the bottom of a coding language blockopen · @AndreasDriesen · 0 Kommentare · 2026-10-06
- #1236 Will zero-copy be implemented on Strix Halo?open · @gmgorag · 1 Kommentare · 2026-10-06
- #1254 [Optimization] Support CPPC / configurable core placement for host loop and expert poolopen · @bceenaeiklmr · 0 Kommentare · 2026-10-06
- #1168 Client-supplied stop sequences are ignored (OpenAI `stop` and Anthropic `stop_sequences`) — blocks agent/harness useclosed · @saadharis · 2 Kommentare · 2026-10-06
- #1239 v0.1.40 two-GPU tester results on a P40 + RTX 3070: 150-request soak clean, --pipeline-windows 2 +9% / +15% decode in 5 of 5 pairsopen · @paulhothersall · 0 Kommentare · 2026-10-06
- #1161 [Feature Request] Port ngram-mod & other ngram-* self-speculative techniques?open · @Interpause · 3 Kommentare · 2026-10-06
- #675 Feature request: `logprobs` / `top_logprobs` on `/v1/chat/completions` (typed decisions from one forward pass)open · @Grandmasg · 14 Kommentare · 2026-10-06
- #1235 Sapphire Rapids (Xeon w7-2475X): +8% decode, +3% prefill and byte-identical greedy output. STRATA_IQ_MT_MIN=1 is faster here, and 8 stager threads beat 32open · @enkynakamura · 0 Kommentare · 2026-10-06
- #1234 Windows/RTX 4090 (0.1.39): stall report stage "decode -1" for 61 s; 0 layers served - thread stacks attachedopen · @skankmoses · 0 Kommentare · 2026-10-06
- #1230 serve/web: add English / Simplified Chinese (zh-CN) UI localizationopen · @btc1000w · 2 Kommentare · 2026-10-06
- #776 --batch on a 100%-resident layer split: engine dies at "captured the batch window over slots 0,1" (zero-doorbell path?)closed · @Seen-Tomorrow · 7 Kommentare · 2026-10-06
- #1211 Bench: --ple-io direct vs. --ple-io ram on system with sufficient headroom of RAMopen · @Thxeverybody · 0 Kommentare · 2026-10-06
- #1212 Review: shared MTP head fix — sibling Q4 state omission and test configuration gapsclosed · @x00r · 1 Kommentare · 2026-10-06
- #1213 .closed · @notGabeNewell · 0 Kommentare · 2026-10-06
- #1200 RTX 5060 Ti 16 GB / Windows / IQ3_S: 36–42 tok/s — tuning directions & method (local gate scan, large pages, PLE FP8)open · @djaafer1975 · 2 Kommentare · 2026-10-06
- #1208 HIP layer split / gfx1030: batched prompt path produces NaN logits (output "!!!!") unless STRATA_HIP_PROMPT_F16=1 and the gfx1030 card is the first stageopen · @cloud4t0r · 0 Kommentare · 2026-10-06
- #1165 Ship an MMQ-enabled engine build (or opt-in loading) for k-quant prompts: a full-RAM machine is GPU-bound on the FP16 dequant pathclosed · @bys13 · 0 Kommentare · 2026-10-06
- #1177 Feature request: add timestamps to engine log linesopen · @Ryutaku · 0 Kommentare · 2026-10-06
- #670 Keeping the old engine when an update succeeds but turns out bad - and a question about where the web app should call itclosed · @demetree · 2 Kommentare · 2026-10-06
- #1155 [Windows / AMD] --vision cpu is still unreachable: hip_vision()'s WIN gate, find_vcvars()'s VS-18 range, and no encoder in the ready-made engineopen · @Hugua700 · 0 Kommentare · 2026-10-06
- #1153 [Volta] STRATA_GDN_CHUNK=1 (PR #1098 chain-split) segfaults the engine on a 4-GPU layer split (4×V100)closed · @justxiami · 1 Kommentare · 2026-10-06
- #1143 Tester report (2× RTX 2080 Ti, Turing): #743 prompt A/B, #776 batch soak, #859 pipeline-windows, #848 resident soak — all passopen · @yusheng227507 · 0 Kommentare · 2026-10-06
- #1141 Windows AMD: prebuilt HIP zip running a model on a discrete card (Radeon AI PRO R9700, gfx1201)open · @KongChengZhi · 0 Kommentare · 2026-10-06
- #1138 Windows: higher process/thread priority improves decode throughput under CPU contention, with a measured background-work trade-offopen · @Unmaple · 0 Kommentare · 2026-10-06
- #881 [Windows / AMD] RX 6800M (gfx1031) runs with a self-built HIP engine — plus 2 Windows bugs found (packaging encoding, vision `vcvars`)closed · @Hugua700 · 3 Kommentare · 2026-10-06
- #1113 sycl: the v0.1.40.1 port does not build as shipped (3 compile errors, 12 undefined symbols)open · @W-IBARI · 0 Kommentare · 2026-10-06
- #975 setup.get_prebuilt_hip() zip-unpack branch returns None on Linux (tools/test_setup_amd.py)closed · @AmjedMVP · 3 Kommentare · 2026-10-06
- #974 Missing --kv-resident in generated args when --resident-budget-gib is set (tools/test_setup_unsloth.py)closed · @AmjedMVP · 3 Kommentare · 2026-10-06
- #977 --check prints [ok] for below-floor RAM and a misleading 'This PC can run Strata' verdictclosed · @AmjedMVP · 2 Kommentare · 2026-10-06
- #973 Golden drift: all 46 test_enter_for_every_question subtests fail on main (tools/test_setup_golden.py)closed · @AmjedMVP · 2 Kommentare · 2026-10-06
- #976 serve: structured-output failure path returns error: null instead of 502 structured_output_failed (test_responses)closed · @AmjedMVP · 3 Kommentare · 2026-10-06
- #1082 Web interface is clitchingopen · @EliyahuAnavim · 1 Kommentare · 2026-10-06
- #1088 Kesebangunan dan Kekogruenanclosed · @yusyennimar-svg · 0 Kommentare · 2026-10-06
- #1085 Performance regression after updating to v0.1.40 - RTX 5090closed · @Predator75 · 1 Kommentare · 2026-10-06
- #1058 0.1.40: a tool call quoted in the thinking is delivered as a real call (fires on a max_tokens cut and inside code fences)closed · @47Hunter47 · 5 Kommentare · 2026-10-06
- #971 Report: UD-IQ4_XS with images works (0.1.39, NVIDIA L40S vGPU in a VM, driver 550 / CUDA 12)closed · @talisp · 0 Kommentare · 2026-10-06
- #1012 0.1.39 serve: after an engine restart, waiting requests hang forever or fail with "list.remove(x): x not in list"closed · @Jackwwg83 · 1 Kommentare · 2026-10-06
- #804 Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`closed · @talisp · 10 Kommentare · 2026-10-06
- #914 strata-vision helper leaks /tmp/strata-vision-* temp dirs (unbounded growth -> OOM on tmpfs /tmp)closed · @cubacrazy · 6 Kommentare · 2026-10-06
- #925 Setup: show progress in the browser from the first second (status page with steps, speed and ETA)open · @storm-ace · 3 Kommentare · 2026-10-06
- #962 Strata ignores cancellation from Claude Codeopen · @awcator · 4 Kommentare · 2026-10-06
- #983 Proposal: Show the elapsed time in seconds during thinking.open · @AndreasDriesen · 1 Kommentare · 2026-10-06
- #952 BrokePipeError when terminating Strata on Linuxclosed · @pdsmike · 1 Kommentare · 2026-10-06
- #874 A lot of disk write when vision is on.closed · @HouZoengLo · 2 Kommentare · 2026-10-06
- #775 Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profileclosed · @JiuYue0820 · 2 Kommentare · 2026-10-06
- #769 server.py: error: port 8080 is already in useopen · @XeonG · 2 Kommentare · 2026-10-06
- #691 Windows: generation slows when the server window is minimized (EcoQoS moves threads to E-cores) — fix includedclosed · @mozophe · 3 Kommentare · 2026-10-06
- #798 Hybrid CPU detection on Linux counts only the Turbo Boost Max 3.0 "favored" cores as P-cores (Core Ultra 7 270K Plus: 2P/22E instead of 8P/16E) — affects setup's `--pool-workers` and `--pool-affinity auto|p-cores`closed · @WUZICANGJIE · 1 Kommentare · 2026-10-06
- #737 [enhancement] #498 follow-up: UD-Q4_K_XL split gate — allow confirm/override just below 135 GB total RAMclosed · @biteric2000 · 2 Kommentare · 2026-10-06
- #730 WARNING: the RAM budget (--resident-budget-gib) cannot be kept - experimental Q4closed · @SladeThe · 2 Kommentare · 2026-10-06
- #767 --vision cpu caps images at 300 tokens; llama.cpp's mtmd asks for >=1024 for Qwen-VL groundingclosed · @miskahm · 3 Kommentare · 2026-10-06
- #830 Let a one-shot request skip the prompt-cache work (~45 ms per call that is never reused)closed · @brenoperucchi · 2 Kommentare · 2026-10-06
- #843 Tool calls silently dropped in agent sessions after an empty assistant turn (fix: skip empty assistant turns when rendering)closed · @jramirez-uab · 2 Kommentare · 2026-10-06
- #795 Main engine strata.exe crashes with STATUS_ILLEGAL_INSTRUCTION (0xC000001D) on non-AVX-512 CPU - 20 WER crashes across v0.1.35/v0.1.38, fault offsets cluster in fixed code regionsopen · @bjfr123321 · 2 Kommentare · 2026-10-06
- #871 Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)closed · @etzerg · 2 Kommentare · 2026-10-06
- #697 Windows HIP / gfx1201: prompt stalls at `GPU rang 0`; post-store `__threadfence_system()` resolves itclosed · @32x-k · 1 Kommentare · 2026-10-06
- #544 Please run several security auditsopen · @bennmann · 5 Kommentare · 2026-10-06
- #1018 [gfx906] gr_up_fast_kernel<GrMulti> aborts the engine with HSA_STATUS_ERROR_MEMORY_APERTURE_VIOLATION on long-context chats (intermittent; the non-fast path is stable)open · @drumblund · 0 Kommentare · 2026-10-06
- #1009 8 GB card: 0.1.39's default prefill is ~40% slower than 0.1.31 (prompt chunk 512 -> 256 when borrowing cache slots)open · @hiru0118 · 0 Kommentare · 2026-10-06
- #577 0.1.38: UD-Q4_K_XL prompts 15-40% slower on a 96 GB PC, from the unbuffered file tier (STRATA_UNBUFFERED_LOAD=0 restores them)closed · @brenoperucchi · 2 Kommentare · 2026-10-05
- #941 UPD: Successfully run on RTX 3090 24gb + Z590E + 64gb ddr4 3200 2chclosed · @gamalsaad7234 · 0 Kommentare · 2026-10-05
- #979 HIP: a gfx1201 hipBLASLt table for 1.2.1 (100201, ROCm 7.2.0) makes the prompt 6-17% slower than no tableclosed · @aswin-dot-R · 1 Kommentare · 2026-10-05
- #821 ROCm on Windows ZIPclosed · @noctrex · 0 Kommentare · 2026-10-05
- #803 Perplexity ~7-10% above llama.cpp on the same GGUF (Qwen3.8-Flash-Next), also at 2K context, every version testedclosed · @mr20399 · 2 Kommentare · 2026-10-05
- #597 setup: a draft vocabulary for French, acceptance 0.52 to 0.64 and decode 11% faster on an RTX 5090closed · @gputier · 4 Kommentare · 2026-10-05
- #940 Question: Pointing to different Chat Templateclosed · @FearL0rd · 1 Kommentare · 2026-10-05
- #518 Feature request: MacOS M-Chip supportclosed · @xyd945 · 5 Kommentare · 2026-10-05
- #573 Can Strata support Intel Arc B580?closed · @xlwang1188 · 6 Kommentare · 2026-10-05
- #900 Add Total Token sent/recived metric in UI dashboardopen · @awcator · 1 Kommentare · 2026-10-05
- #542 sm_120 + IQ packs: ~4x prefill regression 0.1.34-0.1.36 when libcudart's ABI is mismatched - the #420 fits() gate silently disables MMQ (smpbo reads 1)closed · @gravitomagnetic · 3 Kommentare · 2026-10-05
- #840 --calibrate on Windows + AMD (HIP prebuilt 0.1.39) measures decode ~18x slower than the same engine through server.py, so the tuning is meaninglessopen · @The-Dude-2020 · 0 Kommentare · 2026-10-05
- #605 --ple-io direct on rotational storage deadlocks prefill; the watchdog reports it as a generic engine stall (triage discriminator + --ple-io ram fix)closed · @cjmckenna · 2 Kommentare · 2026-10-04
- #738 Configurable system prompt in the web GUIopen · @RobinMeyer · 1 Kommentare · 2026-10-04
- #641 AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64): a working port, numbers, and PRs to upstream itclosed · @JeanP00l · 6 Kommentare · 2026-10-04
- #590 RX 7900 XTX (gfx1100), Ubuntu, setup's ROCm 7.10 wheel: ~0.8 tok/s decode, GPU at 100% / ~320 Wopen · @ibrokers77882 · 3 Kommentare · 2026-10-04
- #649 HIP (gfx1030): intermittent 'verify: timed out at layer N (#267)' with UD-Q4_K_XL — repro dataopen · @jollyroger1480 · 3 Kommentare · 2026-10-04
- #747 Feature request: multi-conversation parking on multi-GPU layer-splitclosed · @AdaxLabs · 1 Kommentare · 2026-10-04
- #621 Please add Unsloth Qwen3.8-Flash-Next-UD-IQ4_XSclosed · @Almazick · 3 Kommentare · 2026-10-04
- #631 Can we prevent the browser from opening to 127.0.0.1 by default?closed · @frederikhors · 3 Kommentare · 2026-10-04
- #585 Windows: source build for sm_70 (Volta) fails to link - LNK1169 (cudart_static.lib vs cudart.lib)closed · @noahark · 1 Kommentare · 2026-10-04
- #623 avx1 not supprted for xeonV2closed · @yotadrivers-blip · 5 Kommentare · 2026-10-04
- #515 Feature request: Intel Arc B390 / Panther Lake iGPU support on Windows (64 GB shared RAM)closed · @JerroldH · 3 Kommentare · 2026-10-04
- #533 Hot VRAM resize: let the engine yield VRAM to other apps and take it backclosed · @AXKore · 3 Kommentare · 2026-10-04
- #587 tools/make_profile.py --base cannot reorder a complete base (shipped profile reports 0 from the traces)closed · @yannickloth · 1 Kommentare · 2026-10-04
- #564 Gui settingsclosed · @XeonG · 2 Kommentare · 2026-10-04
- #596 Add Conversation Cache Management UI to Monitor Tabclosed · @StoyanStAtanasov · 1 Kommentare · 2026-10-04
- #537 Literal </think> quoted in reasoning yields empty stop or leaked reasoning (v0.1.37)closed · @fenrir-labs76 · 6 Kommentare · 2026-10-04
- #609 --no-browser flag?closed · @mkultra333 · 4 Kommentare · 2026-10-04
- #629 Re-running setup with a different --context silently discards mcp_servers, mcp and sampling from the run configclosed · @jmaher-fas · 1 Kommentare · 2026-10-04
- #592 serve: a malformed "tools" value kills the request thread (connection reset / 502 instead of a 400)closed · @MingShi350 · 1 Kommentare · 2026-10-04
- #604 Resident expert placement on a layer split, for mixed or older GPUs: where I think it helps, plus some doc notes, and an offer to testclosed · @paulhothersall · 4 Kommentare · 2026-10-04
- #610 RTX A3000 / 0.1.38: a native IQ3_XXS decode window decomposes via STRATA_VERIFY_PROFILE, and its GPU side is memory-bandwidth-boundclosed · @yannickloth · 1 Kommentare · 2026-10-04
- #588 Reported expert-cache hit rate excludes --pcie-frac experts, so it rises as tok/s fallsclosed · @yannickloth · 1 Kommentare · 2026-10-04
- #601 Linux build fails on Ubuntu 26.04 (glibc 2.43) with CUDA 12.9 (nvcc host-math conflict)closed · @zimuhuan-code · 1 Kommentare · 2026-10-04
- #633 The pinned expert arena is allocated without a host-RAM check: a container limit or a low MemAvailable ends in the OOM killer mid-load, not a message (the file tier already has the probe)closed · @Avicennasis · 1 Kommentare · 2026-10-04
- #644 Specifying explicit GPU split crashes Strataclosed · @tonydiep · 1 Kommentare · 2026-10-04
- #620 IQ3_S: "native head upload: out of memory" at startup when the desktop runs on a second GPU (more free VRAM on the card)closed · @gu-feng · 1 Kommentare · 2026-10-04
- #756 Title: Windows: llama.cpp UI ("llama-ui") opens at http://127.0.0.1:8080/. Where is the Strata web app / Monitor? (engine 0.1.38)closed · @ikura2024 · 2 Kommentare · 2026-10-04
- #505 gfx1100 (W7800 48 GB, 30 GB RAM): full-resident experts — decode 76.5 t/s @ ~100K ctx on 0.1.34 (no thinking), + hipBLASLt 100202 table, IQ2_XS vs IQ3_XXS reversedclosed · @lawsirlawsir-png · 5 Kommentare · 2026-10-04
- #725 Issue: API key provided via --api-key flag is not applied in the UIopen · @engharat · 0 Kommentare · 2026-10-04
- #617 V100-32GB field report: throughput, a thinking-budget pitfall, and Flash-125B vs Qwen3.8-27B on the same boxclosed · @noahark · 3 Kommentare · 2026-10-04
- #602 Measured numbers: Qwen3.8-Flash-Next on 2× RTX 3060 (Q2_0 vs IQ3_S), plus a local reproduction of #75closed · @zimuhuan-code · 2 Kommentare · 2026-10-04
- #684 [Question] Prompt reProcessingclosed · @ukrolelo · 2 Kommentare · 2026-10-03
- #676 Feature request: Multi-chat sidebar in web UIopen · @vyazham · 0 Kommentare · 2026-10-03
- #671 serve: ordered lists whose items are separated by blank lines render every item as "1."open · @kdmcser · 0 Kommentare · 2026-10-03
- #552 Is there any way to continue when the context llimit reached?closed · @SoftologyPro · 5 Kommentare · 2026-10-03
- #543 Can we add OpenCode config.jsonc to docs?closed · @frederikhors · 4 Kommentare · 2026-10-03
- #557 AMD MultiGPUclosed · @Elefant-Freeciv · 5 Kommentare · 2026-10-03
- #647 Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS works on double Vega 20 GPU (Radeon Pro VII 16gb each).closed · @Marczellito · 1 Kommentare · 2026-10-03
- #616 RTX A3000 / 0.1.38: AVX2 i-quant gate/up codebook gather — bit-exact ~1.3x isolated, +1.0% e2e, kept as opt-inclosed · @yannickloth · 1 Kommentare · 2026-10-03
- #574 Please enable support for China's Alipay or WeChat Pay.closed · @hhiew · 1 Kommentare · 2026-10-03
- #607 Offer to donate AI inference and contribution work to Strataclosed · @obselate · 1 Kommentare · 2026-10-03
- #643 May I ask if this can be made into an uncensored version?closed · @mymoses · 7 Kommentare · 2026-10-03
- #534 Performance, RTX3060 vs RX7800XTclosed · @blastbeng · 3 Kommentare · 2026-10-03
- #545 Error: 400: {"type":"invalid_request_error","message":"prompt (38963 tokens) + max tokens (226281) exceeds the context (262144); requests are never truncated"}closed · @avarakin · 4 Kommentare · 2026-10-03
- #516 desktop (KDE/Wayland): compositor VRAM exhaustion when expert-cache auto fills a 24 GB card - pin framebuffer -12, KWin graphics resetclosed · @xyzzing · 4 Kommentare · 2026-10-03
- #520 [2x RTX PRO 6000] Would this approach allow me to run full bf16 model?closed · @lulinxuan · 2 Kommentare · 2026-10-03
- #513 Feature Request: experimental support for Qwen3.8-Whittle-MoE-27B-A17.8Bclosed · @Vitaliy86 · 1 Kommentare · 2026-10-03
- #514 Feature Request: use AMD iGPU instead of CPU for offloadingclosed · @Vitaliy86 · 1 Kommentare · 2026-10-03
- #556 EXL3 (exllamav3 trellis) weight supportclosed · @teee9 · 1 Kommentare · 2026-10-03
- #539 Strata support AMD's RDNA2, specifically the RX 6900 XT.closed · @aidoluiz · 1 Kommentare · 2026-10-03
- #548 decode_cluster_parity: graph replays race their input upload (pageable cudaMemcpy + non-blocking stream)closed · @xenodeve · 2 Kommentare · 2026-10-03
- #560 AMD/Linux desktop: default VRAM reserve (700 MiB) lets the driver evict ~24 GB to RAM, OOM kills kwin; --vram-reserve-mib 3072 fixes itclosed · @icodebot · 1 Kommentare · 2026-10-03
- #522 Helloclosed · @ebdellifakher018-rgb · 1 Kommentare · 2026-10-03
- #511 Engine stopped unexpectedlyclosed · @sinand99 · 1 Kommentare · 2026-10-03
- #530 high/xhigh reasoning_effort can silently run to max_tokens with empty content when reasoning_budget_tokens isn't setclosed · @maxlippe-leonardo · 1 Kommentare · 2026-10-03
- #549 Update.sh fails to update with a KeyError: 'args' errorclosed · @bscout9956 · 1 Kommentare · 2026-10-03
- #551 Wrong UI on Windows?closed · @SoftologyPro · 4 Kommentare · 2026-10-02
- #498 setup: offer the layer split (`--gpus`) for UD-Q4_K_XL when the GGUFs fit in RAM — the engine already runs it without the budget (2x RTX 3090: 31 → 64-78 tok/s)closed · @mad9home · 1 Kommentare · 2026-10-02
- #507 Naive question about context windowclosed · @jefmud · 3 Kommentare · 2026-10-02
- #509 layer_split auto is 2.86x slower than a balanced K on prefill with a heterogeneous pair (3080 + V100, measured)closed · @meelonjisoo-commits · 2 Kommentare · 2026-10-02
- #502 Possible support for Intel cards - B60closed · @Pacoboyd · 1 Kommentare · 2026-10-02
- #506 RTX 5090 (sm_120): prefill ~3x slower on 0.1.34 than 0.1.31 (IQ3_S, single GPU)closed · @gravitomagnetic · 1 Kommentare · 2026-10-02
- #503 No Support For Q4_K_Mclosed · @CrystalNya · 0 Kommentare · 2026-10-02