Pull Requests(引擎)
社区 · 892 条
- #1490 Add a workflow that builds and pushes the Docker image on a version tagopen · @nullata · 0 评论 · 2026-10-08
- #1307 Add Dockerfile for building ROCm-based Strata Docker imagesopen · @jintakhan · 0 评论 · 2026-10-08
- #1486 feat: introduce Dockerfile.rocm-stableopen · @dambaev · 0 评论 · 2026-10-08
- #1497 docs: tool selection and recovery analysis with model-family guideopen · @CC-David-CC · 0 评论 · 2026-10-08
- #1491 bench: completed Flash Next tool-use results with charts and examplesopen · @CC-David-CC · 0 评论 · 2026-10-08
- #423 INTEL B70 32gb supportclosed · @maxfridbe · 0 评论 · 2026-10-08
- #1499 core: write reduced helper rows directly to mapped outputopen · @W1nge · 0 评论 · 2026-10-08
- #1204 serve: a Chinese / English language switch for the web appopen · @hogtinbao · 0 评论 · 2026-10-08
- #1498 KV streaming under WSL: pin the K/V host copy with cudaHostRegisteropen · @yjamil8 · 0 评论 · 2026-10-08
- #1315 glm5-next: GLM-5.3-Flash support {intial attempt & prefill really sucks avoid testing its not ready }open · @gopinath87607 · 0 评论 · 2026-10-08
- #1472 Intel Arc A770 (DG2, 16 GB) support in the SYCL portopen · @samsonnguyen · 0 评论 · 2026-10-08
- #1496 Community benchmark: RTX 3070 Ti 8 GB, Ryzen 9 5900X, 64 GB DDR4-3600 (Windows 11)open · @gbwzzy218 · 0 评论 · 2026-10-08
- #1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMAopen draft · @k93k2J-glitch · 0 评论 · 2026-10-08
- #1492 fix(serve): save and restore sessions without an MTP drafteropen · @Momoyeyu · 0 评论 · 2026-10-08
- #1271 serve: --conversation-cache-spill-dir keeps evicted conversations on disk across restartsopen · @ANBAL534 · 0 评论 · 2026-10-08
- #1493 Experimental cooperative memory relief and resumable text requestsopen draft · @midhatn · 0 评论 · 2026-10-08
- #1362 serve: preserve Markdown state across skipped tool-call newlinesopen · @hawkli-1994 · 0 评论 · 2026-10-08
- #1456 bench: Flash Next variant tool-eval campaignclosed draft · @CC-David-CC · 0 评论 · 2026-10-08
- #1331 Aporte/disk mirror y sesiones agnosticasopen · @shahrokhzargarpour · 0 评论 · 2026-10-08
- #1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)open · @routhjim · 0 评论 · 2026-10-08
- #1489 Responses persistence and disk-only cache: CUDA/HIP integrationopen · @CC-David-CC · 0 评论 · 2026-10-08
- #1038 decode: STRATA_PCIE_BALANCE=1 picks each layer's PCIe count from measured costs (opt-in)closed · @sergqwer · 0 评论 · 2026-10-08
- #1441 Prompt chunks sized for a layer split's pipeline (STRATA_PREFILL_PIPE=1, opt-in)open · @ruibeikaa · 0 评论 · 2026-10-08
- #1415 cpu: add bit-exact AVX2 singleton expert rowsopen · @InB4DevOps · 0 评论 · 2026-10-08
- #1426 dflash: experimental standalone DeepSpec DFlash block drafter for Qwen3.8-Flash-Next (greedy, correctness-first)open draft · @j-luwierski · 0 评论 · 2026-10-08
- #1327 serve: opt-in finite video inputopen · @constantindjonkam · 0 评论 · 2026-10-08
- #1087 cpu: allow configuring expert-pool tasks per phaseclosed · @Unmaple · 0 评论 · 2026-10-08
- #889 sycl: read doorbell flags with an uncached L1+L3 hint (the GPU never saw the host's store on an Arc Pro B60)closed · @joeyjoe02 · 0 评论 · 2026-10-08
- #1481 web: add Copy below code blocks (#1257)open · @BGonnermann · 0 评论 · 2026-10-08
- #1458 helper cache: configurable VRAM allowance with allocation-floor checksopen · @saikiran-rs · 0 评论 · 2026-10-08
- #1459 bench: RX 7900 XTX + 6800 XT helper layout, 1M and long-output resultsopen · @saikiran-rs · 0 评论 · 2026-10-08
- #1424 Pascal GP100: a Q8_0 decode GEMV (STRATA_Q8_SM60=1, opt-in) and IQ3_XXS / IQ3_S tables in shared memoryopen · @ruibeikaa · 0 评论 · 2026-10-08
- #1319 serve: add opt-in bounded Responses history and continuationopen · @mdwsk88 · 0 评论 · 2026-10-08
- #1402 Qwen3.6-35B-A3B and Ornith-1.5-35B-A3B (qwen35moe): a second, smaller model for 8-12 GB cardsopen · @aflin · 0 评论 · 2026-10-08
- #1355 docs(bench): V100 follow-ups: 0.1.40 engine, UD-IQ4_XS tier, context decay, --parallel 4open · @noahark · 0 评论 · 2026-10-08
- #1477 docs: correct pipeline async compatibility (#1447)open · @BGonnermann · 0 评论 · 2026-10-08
- #1228 serve: list items separated by blank lines are one list, and a list continued later keeps its numbers (#671)open · @BGonnermann · 0 评论 · 2026-10-08
- #1395 hip: a hipBLASLt tuning table for gfx1150 (Radeon 890M), and gfx1151's exact speed switches on gfx1150open · @zeriyoshi · 0 评论 · 2026-10-08
- #1390 Windows: native Intel Arc build and setup (OpenCL), plus three Windows-SDK macro fixesopen · @demetree · 0 评论 · 2026-10-08
- #1401 V100 (sm_70): prompt experts on FP16 tensor cores (+10% prompt) and three existing decode kernels as defaults (-6.8% GPU per window)open · @rewin123 · 0 评论 · 2026-10-08
- #1476 Update GLM port to Project Maya v1.0.4open draft · @apr3ndi5 · 0 评论 · 2026-10-08
- #1241 serve: the Responses API reads Codex's additional_tools input items as tools (#782)closed · @BGonnermann · 0 评论 · 2026-10-08
- #1167 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)open · @xjc10 · 0 评论 · 2026-10-08
- #1151 hip: the FP32 tiled QSA block scorer on HIP cards (opt-in STRATA_SELECT_SIMT=1, as on CUDA; a 128K prompt +7.4% on gfx1030)open · @xjc10 · 0 评论 · 2026-10-08
- #1471 feat: optionally offload idle session state while keeping model weights loadedopen · @W1nge · 0 评论 · 2026-10-07
- #1465 cuda: retain HC norm products in existing scratchopen · @W1nge · 0 评论 · 2026-10-07
- #1455 serve: reuse prompt tokens across safe shared boundariesopen · @W1nge · 0 评论 · 2026-10-07
- #1454 prefill: reuse embedding storage for half outputsopen · @W1nge · 0 评论 · 2026-10-07
- #1451 prefill: allocate PLE host buffers only when neededopen · @W1nge · 0 评论 · 2026-10-07
- #1377 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.2open · @Dmitry-B · 0 评论 · 2026-10-07
- #1457 prefill: reuse shared scratch for the HC readopen draft · @W1nge · 0 评论 · 2026-10-07
- #1467 bench: community report - RTX 4090, Flash-Next IQ3_XXS at 204800, engine 0.1.40.3open · @Dmitry-B · 0 评论 · 2026-10-07
- #1462 bench: RTX 5090 Laptop GPU follow-up for 0.1.40.2 and 0.1.40.3open · @wolffahrer · 0 评论 · 2026-10-07
- #1461 Peer tier with mutual help: concurrency 2 at 155 t/s aggregate on dual RTX 3090 (elastic peer pair, opt-in)open · @q8atnight · 0 评论 · 2026-10-07
- #1460 Community benchmark: RTX 5080 in Docker Desktop on Windows (WSL2), Coder IQ1_M, 128K and 64Kopen · @xWinIcex · 0 评论 · 2026-10-07
- #1114 Monitor: per-GPU view, more cards, and a layout you can rearrangeopen · @Efs-O · 0 评论 · 2026-10-07
- #1420 Add optional GLM-5.3-Flash support from Project Mayaopen draft · @apr3ndi5 · 0 评论 · 2026-10-07
- #1452 bench: community report, Tesla V100 32 GB + P100 16 GB expert-helper (mixed Volta/Pascal), Flash-Next IQ3_XXSopen · @aleesposito85 · 0 评论 · 2026-10-07
- #1450 web: wide layout — a four-column single page on very wide screens, and a card look for the chat and the draweropen · @Astemiir · 0 评论 · 2026-10-07
- #1368 prefill: quantize each token's MoE input once and scatter it to its k rows (the same bytes)open · @sergqwer · 0 评论 · 2026-10-07
- #1367 prompt attention: STRATA_PROMPT_ATTN_IMMA=1 - the int8-KV kernel on INT8 tensor cores (opt-in, closer to FP32)open · @sergqwer · 0 评论 · 2026-10-07
- #1313 hipBLASLt table for gfx1201 at 100200 (R9700 calibrated, +15% prefill)closed · @Faks · 0 评论 · 2026-10-07
- #1449 Add Metal backend support for Strata on macOSopen · @dennis-akimov · 0 评论 · 2026-10-07
- #1448 #606 follow-up: clamp the three remaining unclamped q8_1 scale/sum emit sitesopen · @MatthewHines · 0 评论 · 2026-10-07
- #1324 file tier: an LRU part of the RAM budget (--resident-lru-gib), elastic under memory pressure (Windows)open · @malloc32 · 0 评论 · 2026-10-07
- #1164 conversation cache: a shared prefix is copied out of a parked conversation instead of taking it over, so the parked one keeps its historyopen · @BlueKingMuch · 0 评论 · 2026-10-07
- #1376 bench: add RX 7900 XT results for IQ3_S and IQ3_XXSopen · @matrix9neonebuchadnezzar2199-sketch · 0 评论 · 2026-10-07
- #1446 Responses: replay fixes, experimental persistence and optional summariesopen · @CC-David-CC · 0 评论 · 2026-10-07
- #1232 enable --kv-grow with --batchopen · @BlueKingMuch · 0 评论 · 2026-10-07
- #1443 Added benchmarks from my setupopen · @rvannos · 0 评论 · 2026-10-07
- #1253 Enable serial multi-GPU batch MTPopen · @ilumn · 0 评论 · 2026-10-07
- #1242 Isolate concurrent image positions and batch-MTP draft pathsopen · @ilumn · 0 评论 · 2026-10-07
- #1414 prefill: the CPU share (STRATA_PREFILL_CPU_SHARE) up to 3,072-token chunks, from mapped experts, on a layer splitopen · @architectds · 0 评论 · 2026-10-07
- #1439 perf(prefill): cache immutable BF16-to-FP16 weight conversionsopen · @agorevski · 0 评论 · 2026-10-07
- #1438 perf(peer): add opt-in byte-capacity-aware expert placementopen · @agorevski · 0 评论 · 2026-10-07
- #1437 perf(cuda): add opt-in shape-tuned sm75 interleaved verifyopen · @agorevski · 0 评论 · 2026-10-07
- #1436 fix(serve): isolate CPU cores for independent Linux replicasopen · @agorevski · 0 评论 · 2026-10-07
- #1435 docs: report RTX 8000 measurements of existing runtime switchesopen · @agorevski · 0 评论 · 2026-10-07
- #1434 telemetry: GPU load, VRAM, temperature and power for AMD cards on Windows (#1380)open · @gt50 · 0 评论 · 2026-10-07
- #1429 bench: 2x TITAN RTX (sm_75) at Strata v0.1.40.3, IQ3_S at 262Kopen · @shrisha108 · 0 评论 · 2026-10-07
- #1433 community bench: RTX 5090 + RTX PRO 4000 Blackwell, UD-Q4_K_XL helper cache, engine 0.1.40.3open · @sophieandreikina · 0 评论 · 2026-10-07
- #1432 bench: Arc Pro B65 Gen4 results on Strata v0.1.40.2open · @timnevits · 0 评论 · 2026-10-07
- #1430 serve: a tool call written <function= NAME> is the call NAMEopen · @itripn · 0 评论 · 2026-10-07
- #1419 bench: replay token IDs and verified windows in serve decodeopen · @imanu86 · 0 评论 · 2026-10-07
- #1418 cuda: opt-in MMVQ activation reuse across output rowsopen · @imanu86 · 0 评论 · 2026-10-07
- #1417 prefill: scope each split-stage loan estimate to its deviceopen · @imanu86 · 0 评论 · 2026-10-07
- #1117 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in; resubmission of #1030 on the new main)open · @imanu86 · 0 评论 · 2026-10-07
- #854 expert plan: a helper GPU's experts are not part of the PCIe share (--remote-expert-opt decode 40 -> 66 tok/s on 2x RX 6900 XT)closed · @xjc10 · 0 评论 · 2026-10-07
- #1325 file tier (Windows): read each expert from experts.bin and a mirror on a second drive at onceopen · @malloc32 · 0 评论 · 2026-10-07
- #1323 prefill: batched unbuffered stager reads; close the experts.bin view once reads are unbuffered (port of #833)open · @malloc32 · 0 评论 · 2026-10-07
- #1382 docs: index entry for the TITAN RTX community reportclosed · @shrisha108 · 0 评论 · 2026-10-07
- #1427 web: Russian UI localization (dictionary + header language picker)open · @Astemiir · 0 评论 · 2026-10-07
- #1219 setup: a Spanish draft subset (--draft-vocab es); Spanish answers dra…open · @elanonimo832 · 0 评论 · 2026-10-07
- #1288 serve: a long read gives way by what is left to read, not prompt length (#656)open · @Jackwwg83 · 0 评论 · 2026-10-07
- #1421 docs/kubernetes: manifests for one Strata server on a GPU node (kubectl, Kustomize, kapp)open · @blange48 · 0 评论 · 2026-10-07
- #1101 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)open · @uncle-daddy-jp · 0 评论 · 2026-10-07
- #1405 tools: add configurable launches and reuse local model downloadsopen · @1it · 0 评论 · 2026-10-07
- #1107 prefill: gather experts in groups on short prompts tooopen · @brenoperucchi · 0 评论 · 2026-10-07
- #1416 prefill: the CPU share's activations quantized on the GPU, copied down beside the routing sync (builds on #1414)open · @architectds · 0 评论 · 2026-10-07
- #1157 Community benchmark: 2x Tesla P100, Flash-Next IQ3_S, 128k contextclosed · @lechlna000 · 0 评论 · 2026-10-07
- #1411 OLDER_GPUS.md: the P100 row now has a measurement (2x P100, IQ3_S)open · @lechlna000 · 0 评论 · 2026-10-07
- #1408 docs: a short "Contributing a change or report" section in AGENTS.md and a test-requests listopen · @paulhothersall · 0 评论 · 2026-10-07
- #1406 bench: community report, RX 7900 XTX (gfx1100) on WSL2, Coder IQ1_M: 0.1.33 vs 0.1.40, prefill streaming, 16k-256k sweepopen · @theGiallo · 0 评论 · 2026-10-07
- #1400 docs: running Strata behind Open WebUIopen · @fioruccione · 0 评论 · 2026-10-07
- #1399 qsa_select_bench: accuracy floor per sample (fails on an RTX 5090 / MSVC build)open · @sergqwer · 0 评论 · 2026-10-07
- #1398 docs: keeping Codex CLI off the network with Strataopen · @fioruccione · 0 评论 · 2026-10-07
- #1396 hip/gfx906: the tree does not build (q6_k MMQ instance and cudaEventBlockingSync)open · @phoenixclyde · 0 评论 · 2026-10-07
- #1394 serve: evict parked conversations to pass the physical-RAM gate instead of dropping the snapshotopen · @Arthur031221 · 0 评论 · 2026-10-07
- #768 Add Dockerfile for building ROCm-based Strata Docker imagesclosed · @jintakhan · 0 评论 · 2026-10-07
- #1179 batch: with --batch-groups, a group's lowest free slot first (+14-22% at 2-3 clients)open · @blange48 · 0 评论 · 2026-10-07
- #1130 Windows: apply commit limits after allocation failureopen · @SladeThe · 0 评论 · 2026-10-07
- #1391 docs: STRATA_PREFILL_STREAM_MIN=128 in the gfx1151 fast configuration (agent-sized prompt reads on the fused experts: 3.4 -> 2.1 s per turn on Strix Halo)open · @routhjim · 0 评论 · 2026-10-07
- #882 bench: community report, Flash-Next IQ2_XS vs a dense 27B as coding agents (RTX 4090 + 32 GB RAM)closed · @T-Crypt · 0 评论 · 2026-10-07
- #1295 Experimental native Windows path for Intel Arcclosed · @demetree · 0 评论 · 2026-10-07
- #1388 hip: gfx1151 hipBLASLt tuning table for setup's ROCm 7.14.0a20260608 (hipBLASLt 1.4.0)open · @alexdns1 · 0 评论 · 2026-10-07
- #917 bench: community report, Qwen3.8-Flash-Next-GSQ-RCO Q2_0 (Strix Halo/Radeon 8060S gfx1151 + 128 GB VRAM)closed · @DTU-travelpals · 0 评论 · 2026-10-07
- #1381 mtp: the draft layer's experts at Q8_0 and BF16, alongside Q2_0open · @gopinath87607 · 0 评论 · 2026-10-07
- #673 Load the vision encoder lazily with the modelclosed · @ANBAL534 · 0 评论 · 2026-10-07
- #1384 test: fix narrowing byte fixtures in Windows conversation testopen · @midhatn · 0 评论 · 2026-10-07
- #1374 experts: IQ1_S on GPU/AVX-2, IQ1_S+IQ1_M on AVX-512, and a measured PCIe shareopen · @Yxmura · 0 评论 · 2026-10-07
- #1383 generate: say that CUDA_LAUNCH_BLOCKING=1 hangs the verify windows (#1341 #964)open · @ischencheng · 0 评论 · 2026-10-07
- #1202 dashboard: Serve the web page under an optional dashboard pathopen · @gener-vscode · 0 评论 · 2026-10-07
- #1283 prefill: bytes_needed sizes the MoE buffers for the layout init picks (with or without an expert source)open · @sergqwer · 0 评论 · 2026-10-07
- #1282 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)closed · @sergqwer · 0 评论 · 2026-10-07
- #1316 spec: re-probe stale full chains when shorter chains payopen · @hulkbig · 0 评论 · 2026-10-07
- #1379 prefill: STRATA_PREFILL_CPU_SHARE=auto times layers with and without the share and shares only while that is faster (#1282)open · @sergqwer · 0 评论 · 2026-10-07
- #1378 bench: community report, RX 6600 8 GB (gfx1032), EPYC 7232P, IQ3_XXS 128K on Linuxopen · @cryptedx · 0 评论 · 2026-10-07
- #1090 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it againopen · @konijiwa110 · 0 评论 · 2026-10-07
- #1372 DeltaNet recurrence: STRATA_GDN_CHUNKED=1 - the prompt's recurrence in 32-token chunks (opt-in, 1.8x faster, other bits)open · @sergqwer · 0 评论 · 2026-10-07
- #1120 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, opt-in)open · @bsvinay · 0 评论 · 2026-10-07
- #1318 docs(AMD_HIP): my R9700 numbers for hipBLASLt 1.4.1 and STRATA_HIP_WMMAclosed · @l33tm4st3r · 0 评论 · 2026-10-07
- #1274 setup: --thinking / --instruct set the sampling every client gets (#1129)open · @victorgabr · 0 评论 · 2026-10-07
- #1370 prefill: opt-in STRATA_HC_UPMIX=1 on CUDA - the hyper-connection up projection with gr_mix_r as its epilogue (sm_80+)open · @sergqwer · 0 评论 · 2026-10-07
- #1287 feat(vision): support Windows HIP image encodingopen · @Yasei-no-otoko · 0 评论 · 2026-10-07
- #1245 bench: UD-IQ4_XS first NVIDIA measurement (RTX 5090 Laptop, 24 GB)closed · @ManfredCh · 0 评论 · 2026-10-07
- #780 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweepclosed · @1314521gjy · 0 评论 · 2026-10-07
- #1365 Dashboard - metrics logsopen · @skyprince999 · 0 评论 · 2026-10-07
- #1289 hip: the gfx1100 100401 table covers the small-T expert GEMM (prompt +31.9%)closed · @leon-strong · 0 评论 · 2026-10-07
- #1281 sampler: coupled drafts with Gumbel-max picks (STRATA_SPEC_GUMBEL=1, opt-in): +7 points draft acceptance, +9% output at temperature 1.0 on Strix Haloclosed · @routhjim · 0 评论 · 2026-10-07
- #1279 perf(cuda): use exact GP100 VMAD for native DP4A emulationclosed · @bschrib · 0 评论 · 2026-10-07
- #1270 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K on 0.1.40.1 - one card, layer split and expert helper; stock, the two gfx103x switches, and #1150 #1151 #1167closed · @xjc10 · 0 评论 · 2026-10-07
- #1268 --batch-mtp: the first slot admission ends the engine ("unsupported native MMVQ GGML type")closed · @crazyaimachine · 0 评论 · 2026-10-07
- #1264 verify: the batched K/V append writes a streamed K8V4 state's host copy (#1188, #1169)closed · @merbanan · 0 评论 · 2026-10-07
- #1263 Community benchmark: RTX 5090 Laptop GPU (24 GB), Windows 11, Strata 0.1.40.1closed · @wolffahrer · 0 评论 · 2026-10-07
- #1260 tests: order MMQ parity transfers on the compute streamclosed · @mcuygjsliy429 · 0 评论 · 2026-10-07
- #1255 serve: /props total_slots is the engine's batch slotsclosed · @midagedev · 0 评论 · 2026-10-07
- #1249 Pipelined batch: two requests run at half speed (pad rows in group windows, unbalanced slot choice)closed · @crazyaimachine · 0 评论 · 2026-10-07
- #1246 build: vmm.hpp declares size_t via <cstddef> (CUDA 13 headers dropped the transitive include)closed · @anon761 · 0 评论 · 2026-10-07
- #1243 Bench/2026 10 05 community rtx 5080closed · @KalAbaddon · 0 评论 · 2026-10-07
- #1240 Avoid unused MTP branch stream on HIPclosed · @rkcth · 0 评论 · 2026-10-07
- #1227 mcp: the install plan follows setup's AMD rules - Windows AMD and vision=cpu on Linux are planned (#990)closed · @BGonnermann · 0 评论 · 2026-10-07
- #1226 Community benchmark: 2x RTX PRO 4500 Blackwell, Swift IQ3_XXS, engine 0.1.40.1 (solo to 256k, --batch after #776)closed · @qni-live · 0 评论 · 2026-10-07
- #1225 bench: add R9 dual RTX 2080 Ti community resultsclosed · @cezartolvai · 0 评论 · 2026-10-07
- #1222 serve: render a Codex compaction with that conversation's tool prefixclosed · @softbearlolz · 0 评论 · 2026-10-07
- #1221 serve: answer Codex thread-title turns without reading themclosed · @softbearlolz · 0 评论 · 2026-10-07
- #1220 fix #1121: preserve shared draft-head metadata in batch MTP slotsclosed · @x00r · 0 评论 · 2026-10-07
- #1218 Check the ready-made engine against GitHub's SHA-256 before installing itclosed · @demetree · 0 评论 · 2026-10-07
- #1216 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokens (reland)closed · @Zauberio · 0 评论 · 2026-10-07
- #1210 populate '/models' also in lazy modeclosed · @vehystrix · 0 评论 · 2026-10-07
- #1207 fix(serve): keep cached images alive through request preparationclosed · @hulkbig · 0 评论 · 2026-10-07
- #1205 update.sh: a rewritten history is not the user's fault, and the messa…closed · @elanonimo832 · 0 评论 · 2026-10-07
- #1201 verify batch: service the doorbell graph per-window (ar_on), not per-…closed · @wOvAN · 0 评论 · 2026-10-07
- #1199 Community benchmark: RTX 3090 eGPU + 64GB, IQ2_XS vs IQ3_XXSclosed · @lucapug · 0 评论 · 2026-10-07
- #1198 engine: document --spec and --spec-min-p in --helpclosed · @elanonimo832 · 0 评论 · 2026-10-07
- #1197 docs: --calibrate loads the model more than once, and starts it when …closed · @elanonimo832 · 0 评论 · 2026-10-07
- #1196 serve: share the thinking budget with other apps, like max_tokens and…closed · @elanonimo832 · 0 评论 · 2026-10-07
- #1195 serve: a POST /settings body without "defaults" must not clear every …closed · @elanonimo832 · 0 评论 · 2026-10-07
- #1193 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262Kclosed · @shrisha108 · 0 评论 · 2026-10-07
- #1192 Community benchmark: 4x Tesla P100 16 GB (Pascal), IQ3_XXS, CUDA 12 engineclosed · @apollo-mg · 0 评论 · 2026-10-07
- #1185 verify: a window graph that finds no VRAM frees older batch slot layouts' graphs instead of failing (#997)closed · @ischencheng · 0 评论 · 2026-10-07
- #1183 serve: release batch control after terminal native ERRclosed · @hulkbig · 0 评论 · 2026-10-07
- #1182 serve: poll cancellation during quiet prefill waitsclosed · @hulkbig · 0 评论 · 2026-10-07
- #1176 strata_pack.py build --force: unmake the old pack before writing, so a build that stops part-way is an unfinished build, not a pack (#634 follow-up)closed · @Avicennasis · 0 评论 · 2026-10-07
- #1175 serve: a tool's "parameters" / "input_schema" that is not an object is a 400 naming it, not an AttributeError on its first call (#592 follow-up)closed · @Avicennasis · 0 评论 · 2026-10-07
- #1174 expert profile: refuse a header version other than 1, naming both numbers and the fileclosed · @Avicennasis · 0 评论 · 2026-10-07
- #1173 bench: community report, RX 7900 XTX on Windows 11, IQ3_S, engine 0.1.40closed · @nytasuk-commits · 0 评论 · 2026-10-07
- #1172 serve: opt-in recovery of tool calls the model writes next to the template's form ("tool_call_recovery": true)closed · @signalnine · 0 评论 · 2026-10-07
- #1171 docs: minor prefill config change results in faster prefillclosed · @CC-David-CC · 0 评论 · 2026-10-07
- #1160 generate: --layer-split auto rejects a split that cannot startclosed · @fukc-gihtub · 0 评论 · 2026-10-07
- #1159 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer splitclosed · @sarge18 · 0 评论 · 2026-10-07
- #1158 bench: community reports, RTX 4090, IQ3_XXS 204800 and IQ3_S 143360 - 0.1.38 to 0.1.40.1closed · @Dmitry-B · 0 评论 · 2026-10-07
- #1150 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)closed · @xjc10 · 0 评论 · 2026-10-07
- #1148 tests: hip_prefill_wmma_gemm_parity skips on a card without matrix cores instead of failingclosed · @xjc10 · 0 评论 · 2026-10-07
- #1142 hip: accept gfx1010 and gfx1011 (RDNA1) in the arch gateclosed · @AcckiyGerman · 0 评论 · 2026-10-07
- #1140 bench: community report, 2x RX 7900 GRE (gfx1100), layer split, 0.1.40 #848/#859/#880 #776closed · @jase100k · 0 评论 · 2026-10-07
- #1137 Community results: Strata 0.1.40 Q4/Q8, MTP and ngram on RTX PRO 6000closed · @CC-David-CC · 0 评论 · 2026-10-07
- #1134 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workload (re-submission of #466)closed · @gravitomagnetic · 0 评论 · 2026-10-07
- #1132 docs: running Strata behind llama-swapclosed · @christopherrobertbrooks-tech · 0 评论 · 2026-10-07
- #1124 Community results: Strata 0.1.40 Q4/Q8 on RTX PRO 6000closed draft · @CC-David-CC · 0 评论 · 2026-10-07
- #1122 Pipelined windows + async adaptive tier together (--pipeline-windows 2 --adapt-async 1)closed · @Hardin22 · 0 评论 · 2026-10-07
- #1115 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34closed · @Efs-O · 0 评论 · 2026-10-07
- #373 HIP: build with Visual Studio 2026's C++ library (MSVC 14.51+)closed · @BlueKingMuch · 0 评论 · 2026-10-07
- #1110 fix(mtp): --batch-mtp exits on the first admission with a draft-vocab head (slot drafters never copy its type)closed · @imanu86 · 0 评论 · 2026-10-07
- #1109 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)closed · @YamineRL · 0 评论 · 2026-10-07
- #1102 fix: batch host paths honor refreshed residency (deadlocks the engine during a prompt loan)closed · @win10ogod · 0 评论 · 2026-10-07
- #1091 qsa_select: FP32 tiled block scores below sm_80 (+22% 128K / +44% 248K prefill on a 2080 Ti)closed · @konijiwa110 · 0 评论 · 2026-10-07
- #1086 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)closed · @shefowl · 0 评论 · 2026-10-07
- #1083 hip/gfx906: the tree does not build (three STRATA_USE_HIP gates miss STRATA_HIP_GFX906)closed · @xxDoman · 0 评论 · 2026-10-07
- #1081 fix(build): qualify size_t in vmm.hpp (GCC 12 rejects the unqualified name)closed · @aleesposito85 · 0 评论 · 2026-10-07
- #1078 bench: RX 6800M (gfx1031) community report - self-built HIP engine on Windowsclosed · @Hugua700 · 0 评论 · 2026-10-07
- #1076 fix(build): qualify size_t in vmm.hpp (GCC 12 rejects the unqualified name)closed · @aleesposito85 · 0 评论 · 2026-10-07
- #1067 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8closed · @xxDoman · 0 评论 · 2026-10-07
- #1066 hip/gfx906: the 0.1.40 tree does not build (three STRATA_USE_HIP gates miss STRATA_HIP_GFX906)closed · @xxDoman · 0 评论 · 2026-10-07
- #1063 fix(mtp): --batch-mtp exits on the first admission with a draft-vocab head (slot drafters never copy its type)closed · @imanu86 · 0 评论 · 2026-10-07
- #1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer splitclosed · @fioruccione · 0 评论 · 2026-10-07
- #1057 prefill: let Stager threads sleep instead of yield-spinning (Linux + --mmap-experts: prompts ~10x faster on DGX Spark, ~1 core instead of ~20)closed · @uncle-daddy-jp · 0 评论 · 2026-10-07
- #417 bench: community report, RTX PRO 4500 x1/x2 + RTX PRO 4000, Threadripper PRO 3975WX, engine 0.1.31closed · @mlfather · 0 评论 · 2026-10-07
- #1035 serve: share the thinking budget with other apps, like max_tokens and…closed · @elanonimo832 · 0 评论 · 2026-10-07
- #1031 serve: a POST /settings body without "defaults" must not clear every …closed · @elanonimo832 · 0 评论 · 2026-10-07
- #1028 Community benchmark: 2x Tesla P40 22 GB (Pascal, experimental CUDA 12), single card and layer splitclosed · @sarge18 · 0 评论 · 2026-10-07
- #1027 bench: community report, RX 6700 XT 12 GB (gfx1031), Ryzen 7 5700G, I…closed · @GPF · 0 评论 · 2026-10-07
- #1016 Bench: RTX 5060 Ti 16 GB on Ryzen 9 7945HX, IQ3_Sclosed · @avarakin · 0 评论 · 2026-10-07
- #1013 serve: a continued batch request reports its own prompt's reuse; input_tokens never 0closed · @CurtisThe · 0 评论 · 2026-10-07
- #1004 serve: /props total_slots is the engine's batch slotsclosed · @midagedev · 0 评论 · 2026-10-07
- #996 hip: accept gfx1010 and gfx1011 (RDNA1) in the arch gateclosed · @AcckiyGerman · 0 评论 · 2026-10-07
- #995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)closed · @lfontanez · 0 评论 · 2026-10-07
- #993 bench: community report, RTX 5070 Ti + Ryzen 9 5900X, 64 GB DDR4-3600 (IQ2_XS / IQ3_XXS / IQ3_S, six-arm A/B)closed · @obiscr · 0 评论 · 2026-10-07
- #991 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing tracesclosed · @ZhongUncle · 0 评论 · 2026-10-07
- #987 serve: /props reports chat_template_caps so Zed offers tools (#986)closed · @1it · 0 评论 · 2026-10-07
- #982 Added bazzite / fedora / multi gpu supportclosed · @runningman84 · 0 评论 · 2026-10-07
- #981 hip: a rocBLAS solution table for the FP16 prompt GEMMs on gfx103x (the GDN projection 7x below 1,152 tokens; short prompts -10 to -14%)closed · @xjc10 · 0 评论 · 2026-10-07
- #930 cpu: IQ3_S / IQ2_S AVX2 decode without the per-index shift/mask/or (1.07-1.6x per row)closed · @sanastasiou · 0 评论 · 2026-10-07
- #958 #583: the fused layout shrinks the MoE buffers only when every layer can take the fused pathclosed · @Krypto-Whitehat · 0 评论 · 2026-10-07
- #956 serve: an API key with spaces around it or outside Latin-1 works, and a refused key is not called a missing one (#725)closed · @BGonnermann · 0 评论 · 2026-10-07
- #955 bench: Arc Pro B65 Gen4 results on patched v0.1.40closed · @timnevits · 0 评论 · 2026-10-07
- #949 cpu: allow configuring expert-pool tasks per phaseclosed · @Unmaple · 0 评论 · 2026-10-07
- #944 hip: gfx11 WMMA prompt attention and selection scorer, gfx1151 hipBLASLt tableclosed · @nsbradley88 · 0 评论 · 2026-10-07
- #943 Add experimental deepMoE Vulkan backend for DeepSeek text chatclosed · @cklxx · 0 评论 · 2026-10-07
- #936 layer split: overlap the stages' hand-off (STRATA_SPLIT_OVERLAP=1, op…closed · @bsvinay · 0 评论 · 2026-10-07
- #934 Per-layer expert cache: each layer's slots are that layer's blobclosed · @Chromadera · 0 评论 · 2026-10-07
- #927 bench: community report, 2x RX 6900 XT (gfx1030) + Ryzen 5 5600X, IQ3_S at 131K - one card, layer split and expert helper, stock 0.1.39 and with #835/#849/#854closed · @xjc10 · 0 评论 · 2026-10-07
- #924 serve: render Codex's compaction with the conversation's tool prefix (Codex compatibility)closed · @softbearlolz · 0 评论 · 2026-10-07
- #923 serve: answer Codex's thread-title requests without running them (Codex compatibility)closed · @softbearlolz · 0 评论 · 2026-10-07
- #920 docs(hip): GPU_PINNED_MIN_XFER_SIZE=1048576 stopped our gfx1030 mmap-experts stalls (#267, #649)closed · @YamineRL · 0 评论 · 2026-10-07
- #919 hip: Windows support for Radeon gfx1151 APUs (on #895)closed · @storm-ace · 0 评论 · 2026-10-07
- #913 bench: community report, RTX 4060 Ti + Xeon E5-2673 v3 (DDR3, PCIe 3.0 x8), Q2_0 and IQ3_S, 0.1.39 path vs #706 + #764 + a busier tierclosed · @1872183316 · 0 评论 · 2026-10-07
- #912 setup: show the GPU's PCIe link in step 1, warn on fewer lanes than the card hasclosed · @1872183316 · 0 评论 · 2026-10-07
- #911 setup: --inspect, a GGUF's real quantization and whether Strata runs it, from its headers onlyclosed · @1872183316 · 0 评论 · 2026-10-07
- #908 setup: ModelScope as a download source (--source), used when Hugging Face's files do not answerclosed · @1872183316 · 0 评论 · 2026-10-07
- #907 calibrate: measure the adaptive expert tier (--adapt-every / --adapt-swaps / --adapt-decay)closed · @1872183316 · 0 评论 · 2026-10-07
- #904 Verify window: PDL on sm_90+, graph branches, batched Q4_0 KV append and PLE post-ops (bit-identical)closed · @Hardin22 · 0 评论 · 2026-10-07
- #902 Community benchmark: Tesla V100-SXM2-32GB on Windows (CUDA 12.6 / MSVC)closed · @Mitsuasa513 · 0 评论 · 2026-10-07
- #895 hip: support Radeon gfx1151 APUsclosed · @DTU-travelpals · 0 评论 · 2026-10-07
- #894 serve: decode chunked request bodies (fixes empty-body 400 behind some proxies)closed · @oscar-investmatic · 0 评论 · 2026-10-07
- #888 hip: chunk-parallel GDN recurrence (WY form), two parity-gated arms - opt-in STRATA_GDN_WY (measured on gfx1100)closed · @xyzzing · 0 评论 · 2026-10-07
- #883 sm_60: PTX vmad for the __dp4a fallback (bit-exact, ~2.2x in isolation)closed · @willowmerestudios · 0 评论 · 2026-10-07
- #878 # Support sm_60, sm_70, and sm_86 with CUDA 12.6 or 12.8closed · @FearL0rd · 0 评论 · 2026-10-07
- #852 engine: --gpu selects the visible GPUs without environment variablesclosed · @anon761 · 0 评论 · 2026-10-07
- #851 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)closed · @Hardin22 · 0 评论 · 2026-10-07
- #850 bench: community report, Tesla V100 32 GB + RTX 4070 layer split with 16 GB of RAM (IQ2_XS)closed · @christopherrobertbrooks-tech · 0 评论 · 2026-10-07
- #849 HIP gfx103x: PR #540's attention kernel with 8 cells per step and DPP lane exchanges (bit-exact, prompt +6-12% on RX 6900 XT)closed · @xjc10 · 0 评论 · 2026-10-07
- #839 Add experimental Linux gfx900 setup with wave64 validationclosed · @sagentlab · 0 评论 · 2026-10-07
- #834 bench: community report, RTX 4090 + 32 GB RAM at 512K context (IQ2_XS)closed · @T-Crypt · 0 评论 · 2026-10-07
- #832 bench: community report, RTX 5070 Ti + Ryzen 7 9800X3D on Windows 11, IQ3_S, 5K to 164K prompt tokensclosed · @Zauberio · 0 评论 · 2026-10-07
- #829 docs: running Strata behind llama-swapclosed · @christopherrobertbrooks-tech · 0 评论 · 2026-10-07
- #823 bench: community report, Tesla V100 32 GB with 16 GB of RAM (Coder IQ1_M)closed · @christopherrobertbrooks-tech · 0 评论 · 2026-10-07
- #815 bench: community results for an RX 6800 on Windows (0.1.39 against 0.1.38)closed · @did-technomancer · 0 评论 · 2026-10-07
- #813 rtx6000fixclosed · @Fringe210 · 0 评论 · 2026-10-07
- #811 Add an RTX 5090 + Ryzen 7 9700X IQ3_S benchmark on engine 0.1.39closed · @steve8697 · 0 评论 · 2026-10-07
- #800 qsa_prompt_attn: Volta (sm_70) m8n8k4 kernel - keep the hi and lo halves in separate accumulator chainsclosed · @eelgaev · 0 评论 · 2026-10-07
- #799 docs: the expert cache size is a budget, and on Windows an over-sized one pages instead of failingclosed · @1314521gjy · 0 评论 · 2026-10-07
- #797 sycl: enable Arc A770 inference and document measured performanceclosed · @atalantisalt03 · 0 评论 · 2026-10-07
- #793 Pipelined windows over the active slots (+ slot allocator), /metrics for Prometheus (vLLM names) + monitoring kitclosed · @blange48 · 0 评论 · 2026-10-07
- #789 prefill: overlap shared-expert work with routed expert uploadsclosed · @InB4DevOps · 0 评论 · 2026-10-07
- #777 Add an RTX 5090 + Ryzen 7 9700X IQ3_XXS benchmark on engine 0.1.39closed · @steve8697 · 0 评论 · 2026-10-07
- #758 bench: community MI50 (gfx906) — Strata 0.1.38 speed and recallclosed · @xxDoman · 0 评论 · 2026-10-07
- #757 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)closed · @30crows · 0 评论 · 2026-10-07
- #750 docs: describe a targeted ROCr USERPTR reclaim workaroundclosed · @zengyishou · 0 评论 · 2026-10-07
- #745 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightly, serve-path decode (Qwen3.8-Flash-Next IQ3_S)closed · @xyzzing · 0 评论 · 2026-10-07
- #744 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)closed · @xyzzing · 0 评论 · 2026-10-07
- #742 qsa_select: FP32 tiled block scores below sm_80 (+22% 128K / +33% 248K prefill on a 2080 Ti)closed · @konijiwa110 · 0 评论 · 2026-10-07
- #740 Community benchmark: RX 7900 GRE (gfx1100, 16 GB), three Flash-Next quantsclosed draft · @jase100k · 0 评论 · 2026-10-07
- #721 Community benchmark: 2× RTX 3060 12 GB, Qwen3.8-Flash-Next IQ3_S (three prompt sizes, 3 runs each)closed · @zimuhuan-code · 0 评论 · 2026-10-07
- #707 docs: community benchmark row - Tesla V100-PCIE-32GB (sm_70 source build, before/after calibration)closed · @noahark · 0 评论 · 2026-10-07
- #698 Bench: RTX 5090 / 9950X3D (IQ3_S, three-pass A/B) + RTX 2080 Ti / 9900KF (IQ3_XXS, Turing)closed · @CYoung83 · 0 评论 · 2026-10-07
- #674 bench: community report, RTX 3090 Ti + 2x Xeon E5-2699 v3 (HP Z840), IQ3_S at 262Kclosed · @pertain99 · 0 评论 · 2026-10-07
- #665 serve: no automatic --layer-split when the args use --peer-deviceclosed · @fks · 0 评论 · 2026-10-07
- #662 setup: prefer uv for the venv and package installs when uv is installedclosed · @tanzious01 · 0 评论 · 2026-10-07
- #648 native dense: run the architecture guard only on the first shard of a splitclosed · @jollyroger1480 · 0 评论 · 2026-10-07
- #632 expert profile: refuse a header version other than 1, naming both numbers and the fileclosed · @Avicennasis · 0 评论 · 2026-10-07
- #628 bench: report Coder IQ1_M on Threadripper 3990X and RX 6900 XTclosed · @Yasei-no-otoko · 0 评论 · 2026-10-07
- #624 bench: community report — RTX 4090, Ryzen 9 7950X, Flash-Next IQ3_Sclosed · @Dmitry-B · 0 评论 · 2026-10-07
- #618 bench: Windows 11 + Radeon AI PRO R9700 (gfx1201), IQ2_XS, 4K to 248K prompt tokensclosed · @HarukiOtaku · 0 评论 · 2026-10-07
- #615 fix(serve): correct cache usage and timings across reasoning continuationsclosed · @liowald · 0 评论 · 2026-10-07
- #570 setup: a read-only --gguf-dir no longer stops setup when marking checked shardsclosed · @gputier · 0 评论 · 2026-10-07
- #569 serve, setup: an empty API key is refused from the config file and from setup's --api-key (#213)closed · @gputier · 0 评论 · 2026-10-07
- #568 parity tests: ple_parity runs without the missing fixtures, gr_parity and ple_parity wait for their setupclosed · @gputier · 0 评论 · 2026-10-07
- #562 serve: /slots reports n_prompt_tokens, so a llama.cpp-style context m…closed · @Zerschranzer · 0 评论 · 2026-10-07
- #532 tests: prefill_mmq_kquant_test waits for its uploads before the kernelsclosed draft · @eadra · 0 评论 · 2026-10-07
- #521 bench: add community RTX 4000 Ada 2-GPU vs 3-GPU layer-split resultsclosed · @lijieming · 0 评论 · 2026-10-07
- #508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35closed · @R6DJO · 0 评论 · 2026-10-07
- #499 bench: community results, RX 9070 XT 16 GB on Windows 11, ready-made AMD engine 0.1.35, IQ3_Sclosed · @homeofe · 0 评论 · 2026-10-07
- #491 docker: GGUF_DIR, RESIDENT_BUDGET_GIB and KV_STREAMING env vars; document the stop timeoutclosed · @mad9home · 0 评论 · 2026-10-07
- #483 bench: community results, 2x RTX 5060 Ti 16 GB (PCIe gen3), Strata 0.1.34closed · @Efs-O · 0 评论 · 2026-10-07
- #476 native head: report the cudaHostAlloc error, not just "cannot pin N MiB"closed · @yannickloth · 0 评论 · 2026-10-07
- #469 bench: community results, RTX 2080 Ti 11 GB + Threadripper 3960X, IQ3_S at 262K with KV streamingclosed · @homeofe · 0 评论 · 2026-10-07
- #466 bench: community report — RTX 5090, IQ3_S at 1M context (YaRN), agent-style workloadclosed · @gravitomagnetic · 0 评论 · 2026-10-07
- #449 fix: fall back when ROCm development files are missingclosed · @kuishou68 · 0 评论 · 2026-10-07
- #440 bench: community report, RTX 5090 + Ryzen 9 9950X3D, IQ3_S at the full 262K window; prefill auto:32768 A/B; conversation-cache 32 GiB democlosed · @Ambolio · 0 评论 · 2026-10-07
- #433 bench: community report, RTX 5090 + Ryzen 9 5950X (AVX2), UD-Q4_K_XL / IQ3_S / Swift IQ3_XXS on engines 0.1.31-0.1.39closed · @brenoperucchi · 0 评论 · 2026-10-07
- #427 tools/vision: portable by default, setup.py/Dockerfile opt into native (#411 #412 #419 follow-up)closed · @homeofe · 0 评论 · 2026-10-07
- #418 bench: community entry, 2x RTX PRO 4500, Swift IQ3_XXS, engine 0.1.30 and 0.1.36closed · @qni-live · 0 评论 · 2026-10-07
- #416 bench: IQ4_XS on a 64 GB PC - pinned arena vs experts.bin mmap vs the RAM budgetclosed · @pipeob0 · 0 评论 · 2026-10-07
- #404 bench(gfx1100): 0.1.31 vs 0.1.30 prefill/decode, curated entry + raw trialsclosed · @rwkeyes · 0 评论 · 2026-10-07
- #401 Chore(github): Create a .github folderclosed · @Joe-Huber · 0 评论 · 2026-10-07
- #389 bench: community report, 2x TITAN RTX (sm_75), IQ3_S at 262Kclosed · @shrisha108 · 0 评论 · 2026-10-07
- #376 HIP: a failed first configure no longer leaves the kernels at -O0closed · @BlueKingMuch · 0 评论 · 2026-10-07
- #351 serve: preserve context across engine restartclosed · @updalla-apshir · 0 评论 · 2026-10-07
- #319 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)closed · @xyzzing · 0 评论 · 2026-10-07
- #1361 tools/hip: gfx1151 hipBLASLt 100500 tuning table (Windows ready-made …open · @sclsgj · 0 评论 · 2026-10-07
- #1360 gfx1151: STRATA_EXPERT_V2K on by default (UD-Q4_K_XL decode experts: same text, +6.5% greedy / +9.9% sampled output on Strix Halo)open · @routhjim · 0 评论 · 2026-10-07
- #1359 setup: link the engine against the toolkit whose nvcc builds it (CUDAToolkit_ROOT)open · @fuad00 · 0 评论 · 2026-10-07
- #1358 build: vmm.cpp builds with a CUDA toolkit older than 12.5 again (#1071)open · @fuad00 · 0 评论 · 2026-10-07
- #1354 docs: the expert cache is a budget, and on Windows an over-sized one pages instead of failingopen · @1314521gjy · 0 评论 · 2026-10-07
- #1351 Community benchmark: Flash-Next IQ3_S on an RTX 4080 SUPER (Windows), with the expert-pool worker sweepopen · @1314521gjy · 0 评论 · 2026-10-07
- #1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)open · @xyzzing · 0 评论 · 2026-10-07
- #1332 calibrate: the PCIe share sweep reaches 1.0, where no expert goes to the CPU poolopen · @aly8246 · 0 评论 · 2026-10-07
- #1333 bench: dual RTX 5090 production and synthetic Strata 0.1.40 resultsopen · @b1naryagent · 0 评论 · 2026-10-07
- #1345 calibrate: the sweep visits every value in every round instead of onceopen · @aly8246 · 0 评论 · 2026-10-07
- #1342 feat: progressive loading (streaming trunk) baselineclosed · @cybort · 0 评论 · 2026-10-07
- #1343 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layersopen · @rluisr · 0 评论 · 2026-10-07
- #1340 core: retain fixed RAM originals during native async exchangesopen · @CC-David-CC · 0 评论 · 2026-10-07
- #1339 cuda: add an opt-in SM120 Q8 T8 QKV MMQ pathopen · @CC-David-CC · 0 评论 · 2026-10-07
- #1338 cuda: reuse secondary Q8 copies in native async refillsopen · @CC-David-CC · 0 评论 · 2026-10-07
- #1336 cuda: compact traversal for read-only Q8 miss fillsopen · @CC-David-CC · 0 评论 · 2026-10-07
- #1335 docs: supporting QFUSE one-token commit evidence for #1209open · @CC-David-CC · 0 评论 · 2026-10-07
- #1334 Adding support for 890Mopen · @htbmixbox · 0 评论 · 2026-10-07
- #1330 Add community benchmark: 2x Tesla P100 16GBopen · @touchtop · 0 评论 · 2026-10-07
- #1329 Fix batch MTP shared draft head typeopen · @zetlaw · 0 评论 · 2026-10-07
- #1328 serve + engine: constrained decoding for JSON formats; accept tools together with json_schemaopen · @ncphri · 0 评论 · 2026-10-07
- #1186 setup: --draft-vocab es - the English/code subset plus the tokens of Spanish textopen · @txiki739 · 0 评论 · 2026-10-07
- #1326 serve: support IPv6 bind addressesopen · @purrsec · 0 评论 · 2026-10-07
- #1321 bench: publish an IQ2_XS speed report for the Radeon AI PRO R9700open · @alexhegit · 0 评论 · 2026-10-07
- #1125 cpu experts: an AVX2 Q8_K activation quantizer (byte-identical to ggml's)open · @Hardin22 · 0 评论 · 2026-10-07
- #1320 test: qualify existing exact HIP paths on gfx906open · @0FL01 · 0 评论 · 2026-10-07
- #1314 serve: keep unfinished tool calls visible as contentopen · @Arthur031221 · 0 评论 · 2026-10-07
- #1237 perf(expert): the cache fill's stage buffers pinned with cudaHostAllocopen · @ZhongUncle · 0 评论 · 2026-10-07
- #1309 bench: community results for an RTX 4090 Laptop GPU (IQ3_S, 128K and 256K)open · @30crows · 0 评论 · 2026-10-07
- #1305 decode: --lookup-chain 1 and 2 do not change the window cap with the default --suffix-draftopen · @ZackO2o · 0 评论 · 2026-10-07
- #1303 decode: --lookup-chain 1 and 2 do not change the window cap with the default --suffix-draftclosed · @ZackO2o · 0 评论 · 2026-10-07
- #1098 prefill/gdn: the Volta prompt recurrence chain-split (+7-15%) and the opt-in chunked pathopen · @ATIVX928 · 0 评论 · 2026-10-07
- #1290 Fix shared MTP slot initialization and speculative row boundsclosed draft · @jeremiahritchey · 0 评论 · 2026-10-07
- #999 hip: the fused-int8 MoE kernels on gfx1100 rocWMMA - opt-in STRATA_PF_FUSED arm + parity fixture (engagement: parity, REJECT-at-parity recorded)closed · @xyzzing · 0 评论 · 2026-10-07
- #1298 hip: the gfx110x fp16 (rocWMMA) qsa-select scorer arm (STRATA_SELECT_WMMA)open · @xyzzing · 0 评论 · 2026-10-07
- #1296 hip: chunk-parallel GDN recurrence (WY form), two parity-gated arms (STRATA_GDN_WY)open · @xyzzing · 0 评论 · 2026-10-07
- #1293 Community benchmark: RX 7900 XTX 24 GB — ROCm 7.1.1 vs ROCm 10.2 nightlyopen · @xyzzing · 0 评论 · 2026-10-07
- #1292 Docs: AMD RX 7900 XTX (gfx1100) IQ3_S rates row (1K-128K)open · @xyzzing · 0 评论 · 2026-10-07
- #1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K contextopen · @NRK-SH · 0 评论 · 2026-10-07
- #950 prefill: opt-in pointer-list bypass for gate/up gather (STRATA_PF_PTR_MMQ)closed draft · @DuncanBetts · 0 评论 · 2026-10-07
- #1047 Adaptive tier: --adapt-decay outside (0, 1) is refused (inf counts, NaN swap gains)closed · @sergqwer · 0 评论 · 2026-10-07
- #1042 prefill: bytes_needed sizes the MoE buffers for the layout init picks (with or without an expert source)closed · @sergqwer · 0 评论 · 2026-10-07
- #1050 prefill: STRATA_PREFILL_CPU_SHARE=auto hands a small chunk's least-routed experts to the idle CPU pool (opt-in)closed · @sergqwer · 0 评论 · 2026-10-07
- #1284 Adaptive tier: --adapt-decay outside (0, 1) is refused (inf counts, NaN swap gains)open · @sergqwer · 0 评论 · 2026-10-07
- #835 HIP gfx103x: the prompt path's 16-bit GEMMs in FP16 in and out (prompt reads ~2x on RX 6900 XT)closed · @xjc10 · 0 评论 · 2026-10-06
- #859 Layer split: pipelined verify windows for one conversation (--pipeline-windows, opt-in)closed · @Hardin22 · 0 评论 · 2026-10-06
- #1278 HIP: run the K-quant MMQ test on HIP builds; UD-Q4_K_XL measured on an R9700open · @ConnorHaggerty · 0 评论 · 2026-10-06
- #1273 serve: session files work with --peer-deviceopen · @pspranger-throw · 0 评论 · 2026-10-06
- #1251 --peer-device prompt share without P2P: host route, per-token sums, FP16 transfers (2.1-2.5x prompts on a no-P2P pair)open · @guthirry · 0 评论 · 2026-10-06
- #1269 serve: restore a session file without holding its K/V in RAMopen · @pspranger-throw · 0 评论 · 2026-10-06
- #1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoringopen · @tntcannon5000 · 0 评论 · 2026-10-06
- #1162 Verify the downloaded engine against GitHub's SHA-256 (replaces #645)closed · @demetree · 0 评论 · 2026-10-06
- #1266 Add the Update the engine cardopen · @demetree · 0 评论 · 2026-10-06
- #1265 Add the engine updater: verify, stage, all-or-nothing swap, rollbackopen · @demetree · 0 评论 · 2026-10-06
- #1262 STRATA_HC_REQ8: hyper-connection projections requantized to int8 + fp32 scale per 32 at load (rebased #704)open · @merbanan · 0 评论 · 2026-10-06
- #1259 setup: --draft-vocab it - the English/code subset plus the tokens of Italian textopen · @fioruccione · 0 评论 · 2026-10-06
- #1247 prefill: --prefill-help and --prefill-help-frac as the official flags for the idle-stage helperopen · @anon761 · 0 评论 · 2026-10-06
- #1187 gfx906: opt-in Q2_0 gate/up weight reuse without singleton regressionopen · @0FL01 · 0 评论 · 2026-10-06
- #595 No-AVX2 CPUs on a native build; run-script context overrideclosed · @4EverBuilder · 0 评论 · 2026-10-06
- #1111 Intel arc 0.1.40open · @maxfridbe · 0 评论 · 2026-10-06
- #1231 kv-grow: a run that cannot lend cache slots maps the whole window up front instead of writing past its first 16K cellsopen · @BlueKingMuch · 0 评论 · 2026-10-06
- #559 Several conversations at once: batch slots, a pipelined layer split, per-stage dense weightsclosed · @blange48 · 0 评论 · 2026-10-06
- #1215 prefill MoE (SYCL): batch the per-expert GEMMs via oneMKL strided gemm_batch (+43% prefill on Arc Pro B70)closed · @verycaptain · 0 评论 · 2026-10-06
- #1224 serve: the Monitor's AMD GPU readings on Windows through ADLX (load, VRAM, temperature, power)open · @tuandat3019 · 0 评论 · 2026-10-06
- #1223 generate: allow the elastic K/V with --peer-device (VMM ranges keep their memory on the owning GPU)open · @chimpera · 0 评论 · 2026-10-06
- #1190 resident RAM mode with a layer split: the copy keeps each stage's lend regionopen · @evaanp · 0 评论 · 2026-10-06
- #1189 Add assistant prefill support for chat completionsopen · @j-luwierski · 0 评论 · 2026-10-06
- #1181 perf(expert): the verify window's GPU plan in O(n) instead of O(n^2)open · @ZhongUncle · 0 评论 · 2026-10-06
- #1184 AMD HIP: the primary ordinal, and the helper caches beside the resident RAM modeopen · @tuandat3019 · 0 评论 · 2026-10-06
- #1170 test(gr): check the fused multi-token read at T=4 and T=2open · @InB4DevOps · 0 评论 · 2026-10-06
- #275 Add bounded disk persistence for evicted conversationsclosed · @jeremiahritchey · 0 评论 · 2026-10-06
- #1163 conversation cache: a re-parked conversation reserves at most 1/8 of each K/V buffer instead of up to 16 MiB, so fewer get evictedopen · @BlueKingMuch · 0 评论 · 2026-10-06
- #1166 cpu pool: --host-core sibling, the host off the interrupts' logical processor (hybrid CPUs too)open · @Hardin22 · 0 评论 · 2026-10-06
- #885 Perf/duplex transfers onlyclosed · @CC-David-CC · 0 评论 · 2026-10-06
- #645 Update the engine from the web app, cautiouslyclosed · @demetree · 0 评论 · 2026-10-06
- #1156 setup: reach the CPU image encoder on Windows + AMD (#1155, #881)open · @Hugua700 · 0 评论 · 2026-10-06
- #1154 decode: support a three-GPU two-window pipeline on 3 P4sopen · @CC-David-CC · 0 评论 · 2026-10-06
- #787 fix: parse tool calls inside reasoningclosed · @kossum · 0 评论 · 2026-10-06
- #1152 Fix Windows AMD telemetry and add optional Arabic UIopen · @almshary · 0 评论 · 2026-10-06
- #1149 HIP gfx103x FP16 prompt (#835 follow-up): STRATA_F16_RANGE measures what reaches FP16's range, and hip_prefill_gemm covers the FP16-io routeopen · @xjc10 · 0 评论 · 2026-10-06
- #1147 docs: translate README to Italianopen · @perronemirko · 0 评论 · 2026-10-06
- #1144 serve: add opt-in instruction skills to Chatopen draft · @medking82 · 0 评论 · 2026-10-06
- #1127 hip: rebased support for Radeon gfx1151 APUsopen · @DTU-travelpals · 0 评论 · 2026-10-06
- #1108 Add architecture check condition in native_dense.cppopen · @cdb0y511 · 0 评论 · 2026-10-06
- #1136 Add opt-in durable chat history, legacy import and browser compactionopen draft · @medking82 · 0 评论 · 2026-10-06
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)open · @j-luwierski · 0 评论 · 2026-10-06
- #1131 generate: the prompt loan keeps its residency bookkeeping without the token graph (#796 part B)open · @j-luwierski · 0 评论 · 2026-10-06
- #711 kv: --kv k8v4 streams with --kv-residentclosed · @T-Crypt · 0 评论 · 2026-10-06
- #1064 Windows: apply commit limits after allocation failureclosed · @SladeThe · 0 评论 · 2026-10-06
- #1126 Harden server, installer and engine against high-risk bugsopen · @aelnaby · 0 评论 · 2026-10-06
- #1123 HIP RDNA2 (gfx103x): prompt GEMMs as FP32 SGEMMs, a faster default that cannot overflow (RX 6800: 343 -> 534 tok/s)open · @vallicgrr · 0 评论 · 2026-10-06
- #905 Pipelined windows + async adaptive tier together (stacked on #859 and #876)closed · @Hardin22 · 0 评论 · 2026-10-06
- #848 Resident RAM mode with a layer split, --vram-reserve-later-mib (#642)closed · @Hardin22 · 0 评论 · 2026-10-06
- #1119 Add local PDF, Word and Excel attachments with screenshot previewsopen · @medking82 · 0 评论 · 2026-10-06
- #796 generate: one startup VRAM plan before the expert cache is committed (#765)closed · @j-luwierski · 0 评论 · 2026-10-06
- #1112 web: make MCP tool activity readable and collapsibleopen · @medking82 · 0 评论 · 2026-10-06
- #1104 docs: visual community RTX PRO 6000 concurrency results (64K/128K)closed · @CC-David-CC · 0 评论 · 2026-10-06
- #732 Overlap streamed KV uploads with prefill computationclosed draft · @moorwu · 0 评论 · 2026-10-06
- #733 Pipeline CPU expert work without a batch-wide phase barrierclosed draft · @moorwu · 0 评论 · 2026-10-06
- #704 STRATA_HC_Q8: hyper-connection projections as int8 + fp32 scale per 32closed · @merbanan · 0 评论 · 2026-10-06
- #1006 HIP RDNA2 (gfx103x): run the prompt path's 16-bit GEMMs as FP32 SGEMMs (RX 6800: prompts 312 -> 452 tok/s)closed · @vallicgrr · 0 评论 · 2026-10-06
- #1007 hip: the QSA block scores on rocBLAS SGEMM, the solution measured on the card (opt-in; a 128K prompt +2.8% on gfx1030)closed · @xjc10 · 0 评论 · 2026-10-06
- #1106 web: follow streaming Thinking and preserve reading positionopen · @medking82 · 0 评论 · 2026-10-06
- #1105 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vision footprintopen · @gopinath87607 · 0 评论 · 2026-10-06
- #1093 Monitor: Load/Unload split button + idle-unload slider, load busy guard, honest unloadopen · @miskahm · 0 评论 · 2026-10-06
- #1100 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)open · @ATIVX928 · 0 评论 · 2026-10-06
- #1099 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +0.9% e2e)open · @ATIVX928 · 0 评论 · 2026-10-06
- #1097 V100: bit-exact GEMV / MoE operator fusions (experimental build only)open · @ATIVX928 · 0 评论 · 2026-10-06
- #1096 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.2% e2e)open · @ATIVX928 · 0 评论 · 2026-10-06
- #1095 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)open · @ATIVX928 · 0 评论 · 2026-10-06
- #946 launcher: pick a model and its settings in one window (experimental)closed · @pipeob0 · 0 评论 · 2026-10-06
- #945 setup: the low-RAM and KV-streaming decisions as functionsclosed · @pipeob0 · 0 评论 · 2026-10-06
- #1092 launcher: pick a model and its settings in one window (experimental)open · @pipeob0 · 0 评论 · 2026-10-06
- #960 serve: prefix snapshots on disk - a new chat restores its system prompt instead of reading it againclosed · @konijiwa110 · 0 评论 · 2026-10-06
- #1077 feat(arm): build and run Strata on aarch64, including DGX Sparkclosed · @johnlockejrr · 0 评论 · 2026-10-06
- #1089 tools: bench_eviction — the R4.1 sweep: compulsory-miss vs LRU vs LFU-decay vs oracle on real routing tracesopen · @ZhongUncle · 0 评论 · 2026-10-06
- #350 Chat: a context gauge, per-request timings, and conversation compactionclosed · @orangeswim · 0 评论 · 2026-10-06
- #366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Nextclosed draft · @j-luwierski · 0 评论 · 2026-10-06
- #933 CPU trellis kernels (AVX2 / AVX-512) for TQ2_T / TQK6 / TQK7 expert-cache misses (stacked on #928)closed draft · @LaurentZuijdwijk · 0 评论 · 2026-10-06
- #928 Gyro (TQ2_T / TQK6 / TQK7 + Hadamard rotation): native support for the agentionai Qwen3.8-Flash-Next quantsclosed · @LaurentZuijdwijk · 0 评论 · 2026-10-06
- #405 docs: add performance documentation and update repository hygieneclosed · @jase100k · 0 评论 · 2026-10-06
- #1084 hip/gfx906: `__dp4a` falls back to a 4-iteration loop although gfx906 has v_dot4_i32_i8open · @xxDoman · 0 评论 · 2026-10-06
- #428 tests: test_setup_golden passes on Linux (<EXE> only as a whole path part)closed · @homeofe · 0 评论 · 2026-10-06
- #462 diagnostics: dump every verify window's logits, so greedy decode can be bisectedclosed · @constantindjonkam · 0 评论 · 2026-10-06
- #472 tests: test_prebuilt_hip_zip passes on Linux (EXE as on Windows)closed · @homeofe · 0 评论 · 2026-10-06
- #478 prefill: MMQ prompt kernels for Q6_K experts (STRATA_MMQ_KQUANTS)closed draft · @yannickloth · 0 评论 · 2026-10-06
- #500 perf(cpu-pool): eliminate host serialization barrier for intermediate activation quantization with adaptive thresholdclosed · @praveshkhatana · 0 评论 · 2026-10-06
- #525 serve: rescue a tool call stranded in an unclosed thinking spanclosed · @chimpera · 0 评论 · 2026-10-06
- #567 serve: tokenise the prompt from the last shared prefix, not from the startclosed · @gputier · 0 评论 · 2026-10-06
- #572 serve: return unfinished tool call text as contentclosed · @Arthur031221 · 0 评论 · 2026-10-06
- #581 serve: the Monitor's per-GPU cards, each card's free VRAM, and the vi…closed · @gopinath87607 · 0 评论 · 2026-10-06
- #611 cpu: group ARM capacities and honor the pool's reserved host coreclosed · @Suoriks · 0 评论 · 2026-10-06
- #619 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layersclosed · @mw00 · 0 评论 · 2026-10-06
- #652 serve: commit only the tokens a verify window hands outclosed · @anon761 · 0 评论 · 2026-10-06
- #660 strata-vision: flash attention off when the CPU encoder's ggml has AVX-512 (3.3x faster there)closed · @shefowl · 0 评论 · 2026-10-06
- #666 setup, serve: the image encoder's device as its own role (--vision-device)closed · @hireymage · 0 评论 · 2026-10-06
- #680 Add persistent chat history, local model selection, and a Windows desktop clientclosed · @0924haruto12 · 0 评论 · 2026-10-06
- #682 serve: a Chinese / English language switch for the web appclosed · @hogtinbao · 0 评论 · 2026-10-06
- #685 Add optional Windows HIP vision encoder builds and package validationclosed · @sifer · 0 评论 · 2026-10-06
- #686 Add a tiny zero-dependency terminal launcher for Strataclosed · @Ayshinko · 0 评论 · 2026-10-06
- #687 populate '/models' also in lazy modeclosed · @vehystrix · 0 评论 · 2026-10-06
- #712 Fix Windows AMD telemetry and add optional Arabic UIclosed · @almshary · 0 评论 · 2026-10-06
- #718 serve: park conversations on disk, keyed by a client id; requests without one stay in RAMclosed · @evaanp · 0 评论 · 2026-10-06
- #719 Add sandboxed HTML previews and scrollable chat code blocksclosed · @arytek · 0 评论 · 2026-10-06
- #724 Add local PDF, Word and Excel attachments with screenshot previewsclosed · @medking82 · 0 评论 · 2026-10-06
- #726 Add opt-in live cache resizing and desktop resource presetsclosed · @medking82 · 0 评论 · 2026-10-06
- #727 Add opt-in durable chat history, legacy import and browser compactionclosed draft · @medking82 · 0 评论 · 2026-10-06
- #741 prefill: BF16 projections on FP16 tensor cores below sm_80 (+8-17% prefill on a 2080 Ti)closed · @konijiwa110 · 0 评论 · 2026-10-06
- #751 Persist conversation KV state in a bounded disk LRU cacheclosed · @Averyyy · 0 评论 · 2026-10-06
- #753 serve: isolate concurrent request metrics (depends on #559)closed draft · @zengyishou · 0 评论 · 2026-10-06
- #763 serve: enforce OpenAI tool_choice and allow structured output with toolsclosed · @cipgysmo · 0 评论 · 2026-10-06
- #792 batch slots: the zero-doorbell graph (100% VRAM resident) rings no layerclosed · @0xPreDa · 0 评论 · 2026-10-06
- #794 split: the auto placement sees each stage's PCIe linkclosed · @chengshenyangtai · 0 评论 · 2026-10-06
- #801 Update LICENSEclosed · @NirmalyaLenka · 0 评论 · 2026-10-06
- #802 Run IQ1_S (Unsloth UD-IQ1_S) on the GPU and AVX-2, and read the expert file tier fasterclosed · @Yxmura · 0 评论 · 2026-10-06
- #809 Intel arc 0.1.39 + perf fixesclosed · @maxfridbe · 0 评论 · 2026-10-06
- #822 Add assistant prefill support for chat completionsclosed · @j-luwierski · 0 评论 · 2026-10-06
- #824 Qwen3.6 35b a3bclosed · @perronemirko · 0 评论 · 2026-10-06
- #827 docs: translate README to Italianclosed · @perronemirko · 0 评论 · 2026-10-06
- #833 Prefill: batched unbuffered expert reads, close the experts.bin view; opt-in RAM budget with helper cachesclosed · @adonizm · 0 评论 · 2026-10-06
- #844 setup: Unsloth's UD-Q3_K_XL, set up as UD-IQ4_XS (its experts are IQ3_XXS / IQ4_NL - no Q3_K)closed · @architectds · 0 评论 · 2026-10-06
- #853 cuda: specialize hyper-connection up for short verify windowsclosed · @InB4DevOps · 0 评论 · 2026-10-06
- #855 Refactor PageCache to fix a 45× read amplificationclosed · @AncientMystic · 0 评论 · 2026-10-06
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)closed · @agurrrrr · 0 评论 · 2026-10-06
- #877 Add a prometheus metrics endpointclosed · @zeevo · 0 评论 · 2026-10-06
- #896 ple: accept Q5_1 and Q8_0 n-gram tablesclosed · @anon761 · 0 评论 · 2026-10-06
- #898 perf: Q8 resident adaptation (+24.5–28.2% rotation gain on RTX PRO 6000)closed · @CC-David-CC · 0 评论 · 2026-10-06
- #901 Add ARM64 NVIDIA GB10 build and native-pack supportclosed · @githubhjs · 0 评论 · 2026-10-06
- #910 Two-GPU decode: bounded attention merge (bitwise, opt-in) and pipelined PLE read-ahead / early chain (stacked on #905)closed · @Hardin22 · 0 评论 · 2026-10-06
- #947 fix: avoid CPU doorbell waits in fully resident batch decodingclosed · @CC-David-CC · 0 评论 · 2026-10-06
- #963 feat(server): stop unchanged completed tool-call loopsclosed · @a00012025 · 0 评论 · 2026-10-06
- #966 serve: recover tool calls the model writes next to the template's form, and return the rest as contentclosed · @signalnine · 0 评论 · 2026-10-06
- #969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)closed · @CC-David-CC · 0 评论 · 2026-10-06
- #992 security-hardeningclosed · @aelnaby · 0 评论 · 2026-10-06
- #994 Isolate concurrent image positions and validate vision recordsclosed · @ilumn · 0 评论 · 2026-10-06
- #998 serve: an engine that exits after an unread ERR line says why (#997 #890)closed · @ischencheng · 0 评论 · 2026-10-06
- #1000 HIP: gfx1102 (RX 7600) community-validated, with an end-to-end run (follow-up to #192)closed · @nexus2905 · 0 评论 · 2026-10-06
- #1001 generate: wait for the residency table's upload before the next window (#871)closed · @ischencheng · 0 评论 · 2026-10-06
- #1002 arena: STRATA_ARENA_MMAP on Windows (mapped experts.bin, streamed first fill)closed · @mirifiuto135-debug · 0 评论 · 2026-10-06
- #1003 resident RAM mode: release each layer's mapped experts as it is copied (Windows)closed · @xdbxdbx · 0 评论 · 2026-10-06
- #1010 batch: up to 3 MTP drafts per slot inside the 8-row window (2 lanes, code: 71.6 -> 84.3 tok/s)closed · @xdbxdbx · 0 评论 · 2026-10-06
- #1011 kv: one pinned KV pool shared by the sessions (--kv-pool-tokens; 4 lanes x 256K pin 6.24 GiB instead of 15.5)closed · @xdbxdbx · 0 评论 · 2026-10-06
- #1014 vision: downscale large images before CPU encoding (12MP: >8 min → 20 s)closed · @Dylan997S · 0 评论 · 2026-10-06
- #1017 feature: RDNA3 support / Cache-aware routingclosed · @dhoard · 0 评论 · 2026-10-06
- #1019 kv_q8: the gather decodes 8 int8 per thread (V100 experimental build, +11%)closed · @ATIVX928 · 0 评论 · 2026-10-06
- #1020 qsa_prompt_attn: the Volta kernel's latency pipeline (32K int8 12.1 -> 10.2 ms, +1.5% e2e)closed · @ATIVX928 · 0 评论 · 2026-10-06
- #1021 V100: bit-exact GEMV / MoE operator fusions (experimental build only)closed · @ATIVX928 · 0 评论 · 2026-10-06
- #1022 prefill/gdn: the Volta prompt recurrence chain-split (+7-15%) and the opt-in chunked pathclosed · @ATIVX928 · 0 评论 · 2026-10-06
- #1023 layer split: the hand-off over peer access on NVLink (3.3 -> 48 GB/s, +1.0% e2e)closed · @ATIVX928 · 0 评论 · 2026-10-06
- #1024 s2_qpn8: opt-in QPN8 expert GEMV on Volta m8n8k4 (M >= 4)closed · @ATIVX928 · 0 评论 · 2026-10-06
- #1025 fix: unblock all-resident batch windows (batch verify waited for a doorbell the graph never publishes)closed · @win10ogod · 0 评论 · 2026-10-06
- #1030 serve: --elastic, an automatic elastic expert cache with a fixed core (opt-in, CUDA VMM)closed · @imanu86 · 0 评论 · 2026-10-06
- #1032 serve: add opt-in bounded Responses history and continuationclosed · @mdwsk88 · 0 评论 · 2026-10-06
- #1033 prefill: gather GPU-resident experts in groups on short prompts tooclosed · @brenoperucchi · 0 评论 · 2026-10-06
- #1036 docs: --calibrate loads the model more than once, and starts it when …closed · @elanonimo832 · 0 评论 · 2026-10-06
- #1037 engine: document --spec and --spec-min-p in --helpclosed · @elanonimo832 · 0 评论 · 2026-10-06
- #1055 two-PC: a second PC runs a block of layers as a stage of the layer split (--stage-server / --remote-stage, decode only)closed · @mirifiuto135-debug · 0 评论 · 2026-10-06
- #1060 Monitor: per-GPU view, more cards, and a layout you can rearrangeclosed · @Efs-O · 0 评论 · 2026-10-06
- #1068 serve: a tool call from the reasoning only when it ends the turn (#1058)closed · @sergqwer · 0 评论 · 2026-10-06
- #1075 web: follow streaming Thinking and preserve reading positionclosed · @medking82 · 0 评论 · 2026-10-06
- #869 serve: opt-in recovery from repeated reasoning passagesclosed draft · @huangserva · 0 评论 · 2026-10-06
- #743 qsa_select: coalesced streaming top-k for long Turing prompts (+23% 248K prefill)closed · @konijiwa110 · 0 评论 · 2026-10-06
- #897 setup: the disk check counts only what step 6 still writesclosed · @bsvinay · 0 评论 · 2026-10-06
- #860 Unsloth UD-Q6_K_XL: Q8_0 PLE rows and Q6_K gate/up expertsclosed · @DaRealDaHoodie · 0 评论 · 2026-10-06
- #965 serve: pin Claude Code's per-request billing stamp, so the conversation cache reuses agent promptsclosed · @signalnine · 0 评论 · 2026-10-06
- #903 prefill: a prompt reads in equal chunks; a short streaming tail stays (#693, rebased)closed · @Zauberio · 0 评论 · 2026-10-06
- #1049 #606: the fused SwiGLU q8_1 quantizers keep the block's scale and sum finite tooclosed · @sergqwer · 0 评论 · 2026-10-06
- #1048 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installedclosed · @sergqwer · 0 评论 · 2026-10-06
- #1046 serve: a request's combined image file is deleted when the request is refused or never runsclosed · @sergqwer · 0 评论 · 2026-10-06
- #1045 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)closed · @sergqwer · 0 评论 · 2026-10-06
- #1044 serve: image sources - network paths refused, URLs capped and fetched outside the FIFO, no local files for other origins' pagesclosed · @sergqwer · 0 评论 · 2026-10-06
- #1043 generate: residency-table uploads wait for their own copy before the non-blocking streams read the tableclosed · @sergqwer · 0 评论 · 2026-10-06
- #1041 serve: an image in an Anthropic tool_result reaches the encoder (Claude Code's Read of a picture)closed · @sergqwer · 0 评论 · 2026-10-06
- #1040 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)closed · @sergqwer · 0 评论 · 2026-10-06
- #1039 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)closed · @sergqwer · 0 评论 · 2026-10-06
- #409 aarch64 / NVIDIA DGX Spark (GB10): build, run and set up with unified memoryclosed · @eelgaev · 0 评论 · 2026-10-06
- #931 serve: control-token text inside a message stays text (<|im_end|> in a file no longer ends the turn)closed · @asp345 · 0 评论 · 2026-10-06
- #396 chore(git): add standard IDEs to gitignoreclosed · @Joe-Huber · 0 评论 · 2026-10-06
- #972 Fix hard-coded path in sycl/tools/build.shclosed · @subsevenx2001 · 0 评论 · 2026-10-06
- #313 RDNA3 WMMA kernels for the gfx1100 prefill (opt-in, runtime-gated on gfx11; rebased on 0.1.30)closed · @StevenChenSE · 0 评论 · 2026-10-06
- #868 docs: Arc Pro B60 rows and notes, and the B60's PCI id (e211) in setup_intel.pyclosed · @LocalXPU · 0 评论 · 2026-10-06
- #866 sycl: keep cudaStreamQuery's answer in the per-layer ring waitsclosed · @LocalXPU · 0 评论 · 2026-10-06
- #865 ngram: add opt-in Q8_0 PLE table supportclosed · @CC-David-CC · 0 评论 · 2026-10-06
- #586 ple: read the PLE table at full BF16 precision (opt-in, alongside FP8)closed · @constantindjonkam · 0 评论 · 2026-10-06
- #599 native experts: Q4_0 and Q4_1 GPU kernels, Q4_0 PLE tableclosed · @saikiran-rs · 0 评论 · 2026-10-06
- #651 ple: read Q8_0 and Q5_1 n-gram tablesclosed · @anon761 · 0 评论 · 2026-10-06
- #734 Reuse long chat history when the last user message is editedclosed draft · @moorwu · 0 评论 · 2026-10-06
- #887 layer split: score every four-way placement instead of guessing itclosed · @gopinath87607 · 0 评论 · 2026-10-06
- #863 cpu kernels: gathered decode on the cores where it is faster, AVX-VNNI rows, IQ3_S one-token kernelclosed · @Hardin22 · 0 评论 · 2026-10-06
- #693 prefill: auto:16384 tries every 1024 tokens above 8192, and equal chunks - prompts 21-38% faster from 20K on a 16 GB cardclosed · @architectds · 0 评论 · 2026-10-06
- #706 Feat/q2 avx2 bitplaneclosed · @merbanan · 0 评论 · 2026-10-06
- #764 decode: STRATA_ADAPT_LAG - #463's reproducibility without its stallclosed · @merbanan · 0 评论 · 2026-10-06
- #614 serve: optionally checkpoint an existing chunk near the prompt tailclosed · @imanu86 · 0 评论 · 2026-10-06
- #378 Elastic K/V (--kv-grow, opt-in): the K/V takes VRAM as the context grows, the expert cache holds the rest (+68% cache on a 16 GB card at 262K)closed · @sergqwer · 0 评论 · 2026-10-06
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)closed · @maverde73 · 0 评论 · 2026-10-06
- #705 Feat/k8v4 kv streamingclosed · @merbanan · 0 评论 · 2026-10-06
- #783 perf(cuda): fused decode/verify/MTP kernels and graph launch reductions on 0.1.39closed · @stuchapin909 · 0 评论 · 2026-10-06
- #807 cuda: opt-in batched expert uploads to reduce host submission overheadclosed draft · @midhatn · 0 评论 · 2026-10-06
- #708 serve: --mtp is optionalclosed · @merbanan · 0 评论 · 2026-10-06
- #731 Avoid duplicate expert slots across helper and primary GPUsclosed draft · @moorwu · 0 评论 · 2026-10-06
- #864 perf: opt-in resident expert exchange buffer rotationclosed · @CC-David-CC · 0 评论 · 2026-10-06
- #876 Asynchronous adaptive expert tier for the resident RAM mode (--adapt-async 1, opt-in)closed · @Hardin22 · 0 评论 · 2026-10-06
- #880 layer split: the stage weight trim works under `auto` too 4X gain in the prefill in 4 way ringclosed · @gopinath87607 · 0 评论 · 2026-10-06
- #846 Keep MTP drafting in concurrent batch slotsclosed draft · @rkcth · 0 评论 · 2026-10-06
- #845 verify: batch windows honor the all-resident stage (no host doorbells)closed · @Yoshiki-Matsuda · 0 评论 · 2026-10-06
- #818 Windows: opt-in release of mapped expert pages after GPU uploadsclosed · @jordicor · 0 评论 · 2026-10-06
- #358 Start: register the expert arena per layer on a thread ahead of the readers (Windows; part 2 of #285)closed · @sergqwer · 0 评论 · 2026-10-06
- #773 file tier: unbuffered reads on Linux too (O_DIRECT + the kernel's asynchronous reads)closed · @cakescats · 0 评论 · 2026-10-06
- #699 Linux: read ahead at startup (cold start ~920 s → 70 s)closed · @jesdga95 · 0 评论 · 2026-10-06
- #749 Respect Windows commit capacityclosed · @SladeThe · 0 评论 · 2026-10-06
- #480 serve, web: a space-free vision temp dir, lazy vision, and a Chinese UIclosed · @ziphao2010 · 0 评论 · 2026-10-06
- #381 Add HIP image encoder support and Linux AMD GPU telemetryclosed · @Wesleyvdk · 0 评论 · 2026-10-06
- #634 strata_pack.py build: refuse an output directory that already holds a pack, naming the fileclosed · @Avicennasis · 0 评论 · 2026-10-06
- #346 setup: don't list the shared Chat settings file as a model config (KeyError: 'exe')closed · @xidus90 · 0 评论 · 2026-10-06
- #608 setup: a 200K context between the 128K rule and 256K (#406)closed · @MingShi350 · 0 评论 · 2026-10-06
- #470 setup: a download stopped before its rename is finished without a request (no 5-minute 416 retry loop)closed · @homeofe · 0 评论 · 2026-10-06
- #424 fix(setup): drop an engine archive that fails to unpackclosed · @alphastorm · 0 评论 · 2026-10-06
- #837 server: "return_progress": true puts the prompt's PP line on the stream as llama.cpp's prompt_progressclosed · @eliaskg · 0 评论 · 2026-10-06
- #861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)closed · @brenoperucchi · 0 评论 · 2026-10-06
- #555 serve: a request's combined image file is deleted when the request is refused or never runsclosed · @sergqwer · 0 评论 · 2026-10-06
- #554 serve: the vision markers inside a message's text stay text and no longer take a picture's place (#150 for the whole marker)closed · @sergqwer · 0 评论 · 2026-10-06
- #785 serve tests: test_json_schema_text_format's schema failure only where jsonschema is installedclosed · @sergqwer · 0 评论 · 2026-10-06
- #762 serve: extract JSON from prose and code fencesclosed · @cipgysmo · 0 评论 · 2026-10-06
- #812 serve: Codex MCP tools (web search) fail as 'unsupported call': map namespace__name calls back to their namespaceclosed · @robwijnhoven · 0 评论 · 2026-10-06
- #810 serve: accept json_schema roots that are unions of object schemasclosed · @canfieldjuan · 0 评论 · 2026-10-06
- #788 serve: tolerate broken stdin pipes during engine cleanupclosed · @siddolo · 0 评论 · 2026-10-06
- #715 serve: say Connection: close on every responseclosed · @evaanp · 0 评论 · 2026-10-06
- #700 serve: a <tool_call> named in prose before the real call is content, not a malformed callclosed · @T-Crypt · 0 评论 · 2026-10-06
- #790 serve: honour OpenAI tool_choice "required" and a named functionclosed · @aguindehi · 0 评论 · 2026-10-06
- #594 serve: read the request body an answer left unread, so the close is not a resetclosed · @gputier · 0 评论 · 2026-10-06
- #454 serve: honor stop and stop_sequencesclosed · @zuraiz-anjum · 0 评论 · 2026-10-06
- #510 fix(serve): safely handle malformed or partial tool call arguments in historyclosed · @praveshkhatana · 0 评论 · 2026-10-06
- #886 serve: skip completed empty assistant turns in chat historyclosed · @Arthur031221 · 0 评论 · 2026-10-06
- #748 serve: invalidate the engine after fatal verification timeoutsclosed · @zengyishou · 0 评论 · 2026-10-06
- #529 serve: an image in an Anthropic tool_result reaches the encoder (Claude Code's Read of a picture)closed · @sergqwer · 0 评论 · 2026-10-06
- #819 frontend: images inside Anthropic tool_result blocks reach the modelclosed · @christopherrobertbrooks-tech · 0 评论 · 2026-10-06
- #582 Serve tool result imagesclosed · @gopinath87607 · 0 评论 · 2026-10-06
- #553 serve: image sources - network paths refused, URLs capped and fetched outside the FIFO, no local files for other origins' pagesclosed · @sergqwer · 0 评论 · 2026-10-06
- #696 web: saved chats and projects on the server, a file viewer, a coding agent, engine start/stop (#361)closed · @homeofe · 0 评论 · 2026-10-06
- #434 Add optional unified Strata Managerclosed · @Ayshinko · 0 评论 · 2026-10-06
- #752 cache: preserve checkpoints when split allocation fails (depends on #559)closed draft · @zengyishou · 0 评论 · 2026-10-06
- #550 generate: residency-table uploads wait for their own copy before the non-blocking streams read the tableclosed · @sergqwer · 0 评论 · 2026-10-06
- #838 #606: the fused SwiGLU q8_1 quantizers keep the block's scale and sum finite tooclosed · @sergqwer · 0 评论 · 2026-10-06
- #873 Fix first-kernel-launch crash on Windows (ROCm/Adrenalin WDDM startup race)closed · @unimo-chan · 0 评论 · 2026-10-06
- #784 sycl: follow #626 (ThreadAffinity) and #559 (NativeDense::load layer range) so v0.1.39 compilesclosed · @sarashin65 · 0 评论 · 2026-10-06
- #778 iq_kernels: portable aligned shared-memory declaration for gfx1030closed · @jollyroger1480 · 0 评论 · 2026-10-06
- #808 hip/gfx906: cudaFuncSetAttribute must be a function, not a macro (ROCm 7.2.1 build fix)closed · @xxDoman · 0 评论 · 2026-10-06
- #786 hip: RDNA3 (gfx1100) WMMA kernel for the int8-KV QSA prompt attention (opt-in STRATA_HIP_WMMA, gfx12/gfx11)closed · @xyzzing · 0 评论 · 2026-10-06
- #720 tools/hip: Add hipBLASLt tuning table for gfx1100 on ROCm 10.0.0 / hipBLASLt 1.4.1 (100401)closed · @jintakhan · 0 评论 · 2026-10-06
- #695 hip: the tuning table's arch gate accepts the RDNA4 sibling (gfx1200 <-> gfx1201)closed · @bsorensen110 · 0 评论 · 2026-10-06
- #565 tools/hip: add gfx1100 hipBLASLt tuning table for hipBLASLt 1.2.2 (100202)closed · @lawsirlawsir-png · 0 评论 · 2026-10-06
- #755 hip: gfx1100 hipBLASLt 1.5.0 tuning table (ROCm 10.2 nightly) — closes the gfx1100 100500 gapclosed · @xyzzing · 0 评论 · 2026-10-06
- #766 hip: a gfx1100 hipBLASLt tuning table for 100401 (packaged ROCm 10.0.0)closed · @miskahm · 0 评论 · 2026-10-06
- #820 prefill: the dense GGUF projections through MMQ on HIP (opt-in STRATA_DENSE_MMQ=1)closed · @xendak · 0 评论 · 2026-10-06
- #826 verify: the shared-expert stream fork starts off on HIP (fixes the 0.1.38 -> 0.1.39 decode regression, #816)closed · @pugolini · 0 评论 · 2026-10-06
- #646 perf(cuda,decode): zero-doorbell resident verify graph, sub-warp expert packing & shared-mem staging (+39-72% tok/s)closed · @stuchapin909 · 0 评论 · 2026-10-06
- #591 Adaptive tier: --adapt-decay outside (0, 1) is refused (inf counts, NaN swap gains)closed · @sergqwer · 0 评论 · 2026-10-06
- #547 prefill: bytes_needed sizes the MoE buffers for the layout init picks (with or without an expert source)closed · @sergqwer · 0 评论 · 2026-10-06
- #295 Pascal build runs (STRATA_EXPERIMENTAL_SM60)closed · @giostrives · 0 评论 · 2026-10-06
- #988 Resident RAM mode with a layer split: only the experts no stage's cache holds (UD-Q4_K_XL on 2x 3090 with 64 GB)closed · @MirkoMorello · 0 评论 · 2026-10-05
- #970 serve: a tool call written inside the thinking ends it implicitly (#804)closed · @talisp · 0 评论 · 2026-10-05
- #598 refill: improve multi-GPU processing speed ~27% PP improvement on 4× RTX 3060closed · @hash13 · 0 评论 · 2026-10-05
- #540 volta (sm_70) and turing (sm_75): faster prompt reading - BF16 projections on the FP16 tensor cores, and a faster pre-Turing attention kernelclosed · @sskver · 0 评论 · 2026-10-05
- #422 hip: enable experimental Strix Halo gfx1151 source buildsclosed · @Donk119 · 0 评论 · 2026-10-05
- #939 serve: a tool call written inside an unclosed thinking span is an implicit think end (#804)closed · @muriloducatti · 0 评论 · 2026-10-05
- #390 Multi-GPU: a --layer-split stage holds only its own layers' weights, and every card's leftover VRAM spills into a helper expert tier (4-way IQ3_S: 9,269 -> 14,172 expert slots, prefill chunk 2,048 -> 4,096)closed · @gopinath87607 · 0 评论 · 2026-10-05
- #584 Split placement searchclosed · @gopinath87607 · 0 评论 · 2026-10-05
- #583 Prefill chunk ringclosed · @gopinath87607 · 0 评论 · 2026-10-05
- #761 stage no longer waits for the whole split below itclosed · @gopinath87607 · 0 评论 · 2026-10-05
- #580 Weight carve 0138closed · @gopinath87607 · 0 评论 · 2026-10-05
- #689 MTP / coding +20% speedup Experimental/rtxpro q8 buffer ownership closed · @CC-David-CC · 0 评论 · 2026-10-05
- #856 hip: RDNA3 (gfx1100) WMMA kernel for the QSA block-score select/scorer (opt-in STRATA_SELECT_WMMA, gfx12/gfx11)closed · @xyzzing · 0 评论 · 2026-10-04
- #368 QSA select: RDNA3 (gfx1100) WMMA block-scores arm on #337's arch dispatch (opt-in STRATA_SELECT_WMMA=1)closed · @xyzzing · 0 评论 · 2026-10-04
- #453 prefill: the draft layer's batched K/V for a ring too (KV streaming) - prefill +4.6% with --kv-residentclosed · @architectds · 0 评论 · 2026-10-04
- #452 qsa_prompt_attn: Q4_0 KV on tensor cores (mode 4) - long prompts with --kv q4_0 ~19% fasterclosed · @architectds · 0 评论 · 2026-10-04
- #439 prefill: a group's expert gathers in one launch (bit-identical, long prompts +4%)closed · @architectds · 0 评论 · 2026-10-04
- #563 Give the expert cache's VRAM back without unloading the model (#533)closed · @Mirtraxxx · 0 评论 · 2026-10-04
- #622 iq: gather the AVX2 codebook from the grid table (STRATA_IQ256_GATHER)closed · @yannickloth · 0 评论 · 2026-10-04
- #640 STRATA_ARENA_MMAP=1: the expert arena as a read-only mapped file (Linux), for small-RAM machinesclosed · @JeanP00l · 0 评论 · 2026-10-04
- #639 layer split: with explicit split points each GPU loads only its own layers' dense weightsclosed · @JeanP00l · 0 评论 · 2026-10-04
- #638 hip: AMD Instinct MI50 / MI60 / Radeon VII (gfx906, wave64) as an opt-in buildclosed · @JeanP00l · 0 评论 · 2026-10-04
- #805 Lazy vision encoder: per-GPU elastic expert cachesclosed · @jmnargi · 0 评论 · 2026-10-04
- #531 multi-GPU: second GPU as an opt-in expert-cache tier (--peer-device), resubmit of #229 on v0.1.36closed · @q8atnight · 0 评论 · 2026-10-04
- #663 prefill: a layer split's idle card streams a share of a one-chunk prompt's expertsclosed · @anon761 · 0 评论 · 2026-10-04
- #656 serve: cooperative prefill preemption at safe chunk boundaries (--prefill-preempt)closed · @j-luwierski · 0 评论 · 2026-10-04
- #655 sm_75: run the prompt path's BF16 products on the FP16 tensor cores (+15-18 % prefill on RTX 2080 Ti)closed · @rafatxf · 0 评论 · 2026-10-04
- #716 conversation parking with --layer-splitclosed · @evaanp · 0 评论 · 2026-10-04
- #653 serve: park conversations with a layer splitclosed · @anon761 · 0 评论 · 2026-10-04
- #526 Allow conversation parking under --layer-split (owner-routed per-stage capture/restore)closed · @chimpera · 0 评论 · 2026-10-04
- #455 setup: mark the Unsloth family NVIDIA only so far in the AMD menu, and point a mixed PC to its NVIDIA card (#429)closed · @zuraiz-anjum · 0 评论 · 2026-10-04
- #456 Monitor: the conversation cache card (parked conversations from the engine's CACHE lines)closed · @sebastianrcnt · 0 评论 · 2026-10-04
- #484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)closed · @xidus90 · 0 评论 · 2026-10-04
- #759 Add native Responses API support for Codex clientsclosed · @Averyyy · 0 评论 · 2026-10-04
- #345 Add native Responses API support for Codex CLIclosed · @quangvu3 · 0 评论 · 2026-10-04
- #571 Feat/responses api 451closed draft · @CC-David-CC · 0 评论 · 2026-10-04
- #664 setup, tools: a draft vocabulary for French, and draft_vocab.py builds one from a text (#597)closed · @gputier · 0 评论 · 2026-10-04
- #683 setup: a re-run of setup carries over the hand-edited config blocks (#629)closed · @hireymage · 0 评论 · 2026-10-04
- #635 serve: a malformed "tools" value is a 400 naming the field, not a closed connection (#592)closed · @Avicennasis · 0 评论 · 2026-10-04
- #701 serve: a malformed "tools" is a 400 naming the field (#592)closed · @bsorensen110 · 0 评论 · 2026-10-04
- #681 serve: a malformed "tools" is a 400, not a dead request threadclosed · @hireymage · 0 评论 · 2026-10-04
- #717 fix: make the CUDA language link the shared runtime (Windows sm_70 LNK1169, #585)closed · @noahark · 0 评论 · 2026-10-04
- #391 iq_avx2: build the IQ2 sign table without a variable shift (BMI2 \shlx\ traps on pre-Haswell CPUs)closed · @demetree · 0 评论 · 2026-10-04
- #394 AVX1: run on a CPU with AVX but no AVX2, with an end-to-end run on Sandy Bridge-Eclosed · @demetree · 0 评论 · 2026-10-04
- #667 Add Intel XPU backend: Qwen decode on Arc Pro B60 via SYCLclosed · @the-talos · 0 评论 · 2026-10-04
- #360 qsa_prompt_attn: Volta (sm_70) crashes in the prompt path - compare the full compute capabilityclosed · @DingoOz · 0 评论 · 2026-10-04
- #600 qsa_prompt_attn: tensor-core kernel for Volta (sm_70, mma.m8n8k4)closed · @fks · 0 评论 · 2026-10-04
- #627 opt for sm_70(just test V100-16g*1 or 2)closed · @ATIVX928 · 0 评论 · 2026-10-04
- #395 Feature/nvidia p40closed · @rafal-prasal · 0 评论 · 2026-10-04
- #524 Adding experimental support for gfx1031 - 6700 XTclosed · @matthewjt · 0 评论 · 2026-10-04
- #442 hip: add community gfx1012 support with version-gated legacy compatibilityclosed · @CC-David-CC · 0 评论 · 2026-10-04
- #677 bench: Linux + 2x AMD Instinct MI50 16 GB (gfx906), Coder IQ1_M, 4K to 128K prompt tokensclosed · @JeanP00l · 0 评论 · 2026-10-04
- #626 cpu: preserve full thread affinity on Windows and Linuxclosed · @Yasei-no-otoko · 0 评论 · 2026-10-04
- #650 pinned: back the Linux expert arena with transparent huge pagesclosed · @anon761 · 0 评论 · 2026-10-04
- #575 qsa select: split the top-k past the register fit - 9% faster long-prompt prefillclosed · @orangeswim · 0 评论 · 2026-10-04
- #603 qsa_select: a 1,024-thread top-k past the register kernel's reach (CUDA) - a 243K-token prompt +26% on an RTX 3060closed · @asp345 · 0 评论 · 2026-10-04
- #578 Improve multi-GPU decode performance via optimized expert-helper executionclosed · @zhedang · 0 评论 · 2026-10-04
- #492 Heterogeneous multi-GPU roles: whole model on one GPU, the MTP draft head on the otherclosed · @hireymage · 0 评论 · 2026-10-04
- #589 make_profile: --reorder so a trace can rank ahead of a complete baseclosed · @yannickloth · 0 评论 · 2026-10-04
- #637 serve: retry an engine start that fails, and keep the last known context after a failed restartclosed · @JeanP00l · 0 评论 · 2026-10-04
- #636 native: add opt-in GBNF constraints to existing generationclosed · @CC-David-CC · 0 评论 · 2026-10-04
- #630 serve: add opt-in stateless Responses APIclosed · @CC-David-CC · 0 评论 · 2026-10-04
- #722 Add Hadamard-INT2 GGUF supportclosed · @SaibaWaipu · 0 评论 · 2026-10-04
- #12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…closed · @chynggi · 0 评论 · 2026-10-04
- #504 web: saved chats in a sidebar, kept in localStorage (#361)closed · @homeofe · 0 评论 · 2026-10-04
- #688 Bench: RTX 5090 / Ryzen 9 9950X3D, IQ3_S, v0.1.34 vs v0.1.38 vs fused prompt kernelsclosed · @CYoung83 · 0 评论 · 2026-10-03
- #355 serve: send Cache-Control: no-cache for index.htmlclosed · @dan64 · 0 评论 · 2026-10-03
- #659 Fix WDDM host pinning for gfx1031 native inferenceclosed · @AxeSl · 0 评论 · 2026-10-03
- #593 prefill: run BF16 GEMMs through FP16 tensor cores on sm_70/sm_75closed · @fks · 0 评论 · 2026-10-03
- #463 decode: wait for an adaptive expert swap before reading the residency tableclosed · @constantindjonkam · 0 评论 · 2026-10-03
- #488 pinned: honor STRATA_NO_LARGEPAGES on Linux; name the hugetlb pool shortfallclosed · @yannickloth · 0 评论 · 2026-10-03
- #487 PCIe probe: median of several primed bursts instead of oneclosed · @yannickloth · 0 评论 · 2026-10-03
- #489 bench/results: RTX A3000 12GB laptop + i7-12850HX (sm_86): hugepages, PCIe probe, calibrationclosed · @yannickloth · 0 评论 · 2026-10-03
- #407 Adaptive tier: --adapt-decay, and --adapt-tuned (opt-in): 31% fewer misses, -19% CPU pool, -22% PCIe on UD-Q4_K_XLclosed · @sergqwer · 0 评论 · 2026-10-03
- #353 NVFP4 routed experts on 0.1.35 (experimental): 120a only for the W4A4 unit (rebase of #292)closed · @sergqwer · 0 评论 · 2026-10-03
- #279 Windows: leave 1500 MiB of VRAM unused by default (decode stalled with 700)closed · @sergqwer · 0 评论 · 2026-10-03
- #546 prefill: the fused layout's smaller MoE buffers only when every layer takes the fused path (STRATA_PF_FUSED=1 on IQ2_XS wrote past them)closed · @sergqwer · 0 评论 · 2026-10-03
- #538 Community benchmark: RTX 5060 Ti 16 GB + EPYC 7B12, UD-Q4_K_XL at 262,144 tokensclosed · @QilinWan · 0 评论 · 2026-10-03
- #441 Contrib/non mtp servingclosed draft · @CC-David-CC · 0 评论 · 2026-10-03
- #576 multi-GPU: conversation parking across a layer splitclosed · @saikiran-rs · 0 评论 · 2026-10-03
- #443 kernels: add opt-in small-window GR and one-warp MMVQ experimentsclosed · @CC-David-CC · 0 评论 · 2026-10-03
- #464 ple: read the n-gram table at full BF16 precisionclosed · @constantindjonkam · 0 评论 · 2026-10-03
- #561 fix: --dump-logits must not promise rows it will never writeclosed · @constantindjonkam · 0 评论 · 2026-10-03
- #536 decode_cluster_parity: the graph replays wait for their uploads (the test failed now and then)closed · @sergqwer · 0 评论 · 2026-10-03
- #527 Experimental P100: pinned RAM complement for fixed layer-split CPU/GPU inferenceclosed draft · @Karenax · 0 评论 · 2026-10-03
- #482 Networked Pool supportclosed · @fragtion · 0 评论 · 2026-10-03
- #382 HIP: MTP prompt pass per group by default (fixes prompt hang on gfx1201 with KV streaming)closed · @abhinand5 · 0 评论 · 2026-10-03
- #387 AMD: ship the gfx1200 hipBLASLt table, with the with/without measurementsclosed · @Efeisot · 0 评论 · 2026-10-03
- #386 tools/hip: gfx1201 hipBLASLt table for ROCm 7.2.4 (hipBLASLt 1.2.2), ~1.8x prompt speed on the R9700closed · @abhinand5 · 0 评论 · 2026-10-03
- #362 File tier: read the GGUF in place unbuffered when the file cache cannot keep it beside the RAM budget (Windows; #286 ported into the existing tier)closed · @sergqwer · 0 评论 · 2026-10-03
- #357 Start: read the expert arena unbuffered when the file cache cannot help (Windows; part 1 of #285)closed · @sergqwer · 0 评论 · 2026-10-03
- #415 IQ4_XS expert rows on the AVX-2 path: the multi-token kernel this format never hadclosed · @pipeob0 · 0 评论 · 2026-10-03
- #363 experts: smaller and fewer launches for the verify window's PCIe callclosed · @BlueKingMuch · 0 评论 · 2026-10-03
- #512 prefill: use active QSA top-k bounds on Turing with large KV capacityclosed draft · @imanu86 · 0 评论 · 2026-10-03
- #473 native experts: Q5_0 GPU kernel (#222)closed · @Rhonstin · 0 评论 · 2026-10-03
- #413 prefill: DeltaNet recurrence with the three value heads of a key head in one thread (bitwise; 1.4x on a 4080 SUPER, 1.3x on a 3090)closed · @BlueKingMuch · 0 评论 · 2026-10-03
- #385 Prompt path: a stager buffer's first job of a generation waits for the previous DMA from it (a narrow race with unpinned blobs)closed · @sergqwer · 0 评论 · 2026-10-03
- #374 Prompt path: the first chunk's n-gram (PLE) rows read beside layer 0, 256 at a time (same output)closed · @sergqwer · 0 评论 · 2026-10-03
- #372 Prompt path: an MMQ group's experts gathered in one launch, after one wait, released by one event (-9% on a 32K prompt, same output)closed · @sergqwer · 0 评论 · 2026-10-03
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cacheclosed · @jeremiahritchey · 0 评论 · 2026-10-02
- #450 calibrate: a gpu list is not a layer split (#447)closed · @EdderTalmor · 0 评论 · 2026-10-02
- #377 HIP: time the PCIe probe on the host on Windowsclosed · @BlueKingMuch · 0 评论 · 2026-10-02
- #380 HIP: count the desktop's VRAM on Windows (WDDM budget, STRATA_WDDM_BUDGET)closed · @BlueKingMuch · 0 评论 · 2026-10-02
- #229 multi-GPU: second GPU as an opt-in expert-cache tier with P2P rows (--peer-device), part 1 of #204closed · @q8atnight · 0 评论 · 2026-10-02
- #309 serve: keep retained K/V through the cache capacity gateclosed · @chimpera · 0 评论 · 2026-10-02
- #523 serve: a tool call that opens inside thinking is an implicit think endclosed draft · @chimpera · 0 评论 · 2026-10-02
- #426 Windows AMD: the HIP backend detects, builds and runs (RX 9070 XT / gfx1201)closed · @RenZekta · 0 评论 · 2026-10-02
- #107 serve: report cached prompt tokens, llama.cpp timings and GET /v1/statusclosed · @architectds · 0 评论 · 2026-10-02
- #134 setup: count the 3-bit models' 262K context against RAM, not a fixed 90 GBclosed · @architectds · 0 评论 · 2026-10-02
- #136 Linux prebuilt from CI: CUDA 12.8, sm_80/86/89 + PTX, attached to each releaseclosed · @architectds · 0 评论 · 2026-10-02
- #501 verify: pin the serve host thread the session loop already pinsclosed · @yannickloth · 0 评论 · 2026-10-02
- #436 serve: a non-streaming request stops when its client disconnects (#430, part of #431)closed · @homeofe · 0 评论 · 2026-10-02
- #435 setup: resuming a download asks only for the missing space, and a finished .part is not re-requested (#425)closed · @homeofe · 0 评论 · 2026-10-02
- #356 HIP: build the AMD engine on Windowsclosed · @jagsan-cyber · 0 评论 · 2026-10-02
- #84 EXPERIMENTAL Rope scaling: contexts past the trained 262K (none / linear / YaRN)closed · @j-luwierski · 0 评论 · 2026-10-01
- #399 fix(setup): drop a refused engine archive instead of reusing itclosed · @alphastorm · 0 评论 · 2026-10-01
- #379 gr_parity: compare the split read (STRATA_GR_V3=1) within float rounding (the test fails with it on)closed · @sergqwer · 0 评论 · 2026-10-01
- #398 setup: stop on EOF instead of accepting prompt defaultsclosed · @Bortlesboat · 0 评论 · 2026-10-01
- #207 E-2 on the AVX-2 path too: prefetch the i-quant expert rows (same switch, CPUs without AVX-512)closed · @pipeob0 · 0 评论 · 2026-10-01
- #42 pinned: MEM_LARGE_PAGES on Windows never worked - enable the privilege, round the sizeclosed · @pipeob0 · 0 评论 · 2026-10-01
- #43 kernels: AVX2 multi-token i-quant row kernels (an AVX2-only CPU decoded every i-quant row once per token)closed · @pipeob0 · 0 评论 · 2026-10-01
- #44 generate: probe the real host->device bandwidth for pcie_fracclosed · @pipeob0 · 0 评论 · 2026-10-01
- #64 parity tool: a new-kernel mismatch must fail the run; dispatch comment fallback was staleclosed · @pipeob0 · 0 评论 · 2026-10-01
- #311 HIP: support RDNA2 gfx1030 working and gfx103X is untested.closed · @xendak · 0 评论 · 2026-10-01
- #339 hip: hipBLASLt tuning table for gfx1201 (R9700) with ROCm 10.2.0a nightly (hipBLASLt 1.5.0)closed · @bsorensen110 · 0 评论 · 2026-10-01
- #337 hip: gfx12 (RDNA4) QSA select - WMMA block scorer, prompt-aware top-k dispatchclosed · @bsorensen110 · 0 评论 · 2026-10-01
- #329 hip: gfx12 (RDNA4) WMMA kernel for the int8-KV QSA prompt attention (1.26-1.35x prompt processing)closed · @bsorensen110 · 0 评论 · 2026-10-01
- #343 Don't pursue: BF16 decode GEMV via Turing tensor cores (sm_75) - measured negativeclosed · @hireymage · 0 评论 · 2026-10-01
- #296 Add OrcaRouter Q4_K_S support (0.1.30)closed · @Suoriks · 0 评论 · 2026-10-01
- #272 feat(cpu): add hybrid architecture awareness and --pool-affinity for P/E-core CPUsclosed · @praveshkhatana · 0 评论 · 2026-10-01
- #291 PLE: the n-gram table in FP8 E4M3, as Qwen ships it (opt-in)closed · @sergqwer · 0 评论 · 2026-10-01
- #290 Token embedding in BF16 from the checkpoint (--embd-gguf)closed · @sergqwer · 0 评论 · 2026-10-01
- #293 KV: optional Hadamard rotation for int8 K/V (STRATA_KV_ROT=1), with the drafter and batched verify made rotation-consistentclosed · @sergqwer · 0 评论 · 2026-10-01
- #280 RoPE: every kernel reads the session's float64 angle tableclosed · @sergqwer · 0 评论 · 2026-10-01
- #283 Prompt path: BF16 projections get the activation's BF16 remainder tooclosed · @sergqwer · 0 评论 · 2026-10-01
- #282 prefill auto: chunks up to 32768 (was 8192), +15% on a 32K promptclosed · @sergqwer · 0 评论 · 2026-10-01
- #281 numerics: saturate the prompt path's FP16 SwiGLU products; log1pf in the unfused GDN gateclosed · @sergqwer · 0 评论 · 2026-10-01
- #317 ple: keep the SSD awake while rows are read (STRATA_SSD_KEEPALIVE)closed · @BlueKingMuch · 0 评论 · 2026-10-01
- #315 hc: the multi-token hyper-connection read split finerclosed · @BlueKingMuch · 0 评论 · 2026-10-01
- #284 Verify: the commit graph no longer waits; it overlaps the MTP draftclosed · @sergqwer · 0 评论 · 2026-10-01
- #332 serve: add a standalone API request monitorclosed · @KadoBOT · 0 评论 · 2026-10-01
- #334 serve: enforce structured JSON in the native samplerclosed · @KadoBOT · 0 评论 · 2026-10-01
- #333 serve: add lazy startup and reliable Windows model cleanupclosed · @KadoBOT · 0 评论 · 2026-10-01
- #321 serve: add CORS preflight (OPTIONS), reverse-proxy header support, and X-Accel-Buffering for SSEclosed · @BlitzenCats · 0 评论 · 2026-10-01
- #320 verify: fix STRATA_VERIFY_PROFILE stage columns (unsigned stamp wrap + slot collision)closed · @xyzzing · 0 评论 · 2026-10-01
- #277 iq_pack: keep experts.bin only when its blobs match the GGUFclosed · @sergqwer · 0 评论 · 2026-10-01
- #287 setup: --draft-vocab cyrillic (English/code + the whole Cyrillic script)closed · @sergqwer · 0 评论 · 2026-10-01
- #288 strata-vision: --flash-attn switch, CPU runs without a CUDA context or warm-up; relative vision pathsclosed · @sergqwer · 0 评论 · 2026-10-01
- #278 serve: count_tokens, Anthropic thinking only when asked, STRATA_REQUEST_LINES, a relative execlosed · @sergqwer · 0 评论 · 2026-10-01
- #274 feat(mtp): add --coupled-draft and --no-coupled-draft CLI options for coupled speculative samplingclosed · @praveshkhatana · 0 评论 · 2026-10-01
- #289 Start: name the GPU and stop at once when the build has no code for it; STRATA_EMULATE_CCclosed · @sergqwer · 0 评论 · 2026-10-01
- #276 diagnostics: STRATA_DUMP_FIRST_LOGITS writes the first generated token's logitsclosed · @sergqwer · 0 评论 · 2026-10-01
- #324 fix(setup): install the engine, model files and packages the checkout was tested withclosed · @alphastorm · 0 评论 · 2026-10-01
- #286 Low RAM on Windows: a pinned tier of the most-read experts, the rest read unbuffered from experts.binclosed · @sergqwer · 0 评论 · 2026-10-01
- #285 Start: read the expert arena unbuffered (the GGUF too, fixes #230), register it while it loads, on its own threadclosed · @sergqwer · 0 评论 · 2026-10-01
- #292 NVFP4 routed experts: converter, pack, decode, prompt path (W4A8, W4A4 on Blackwell), AVX-512 CPU rowsclosed · @sergqwer · 0 评论 · 2026-10-01
- #294 Layer split auto: extract the planner into layer_split.hpp, pin the whole arena outside Windowsclosed · @giostrives · 0 评论 · 2026-10-01
- #314 feat: Qwen3.8-Flash-Next dense model, native Qwen3.5 attention kernel, strata-dense CLIclosed · @perronemirko · 0 评论 · 2026-10-01
- #318 iq_pack: store an F32 ple_conv1d as F16closed · @cripto-bot · 0 评论 · 2026-10-01
- #330 V100 (sm_70): the compute-capability floor drops from 7.5 to 7.0closed · @taweili · 0 评论 · 2026-10-01
- #336 cuda: add experimental Tesla P4 8GB (sm_61) supportclosed draft · @CC-David-CC · 0 评论 · 2026-10-01
- #323 hip: add experimental RX 5500 XT 8 GB (gfx1012) supportclosed · @CC-David-CC · 0 评论 · 2026-10-01
- #325 AMD working on Windows (9070XT tested)closed · @dvasdekis · 0 评论 · 2026-10-01
- #302 HIP/Windows: batch expert copy dependencies in long prefillclosed · @araujoluks · 0 评论 · 2026-10-01
- #322 hip: fast packed-byte intrinsics for RDNA3/RDNA4 (v_perm_b32 + SWAR)closed · @bsorensen110 · 0 评论 · 2026-10-01
- #312 QSA select: ROCWMMA block-scores arm (opt-in STRATA_SELECT_WMMA=1, gfx1100)closed · @xyzzing · 0 评论 · 2026-10-01
- #270 qsa_prompt_attn: run the f16 tensor-core path on Turing (sm_75)closed · @kenh0u · 0 评论 · 2026-10-01
- #257 q2_avx2: two-block 256-bit unpack for the AVX2 Q2_0 expert rows (bit-exact)closed · @hireymage · 0 评论 · 2026-10-01
- #256 HIP: support RDNA4 gfx1200 (RX 9060 XT)closed · @Efeisot · 0 评论 · 2026-10-01
- #254 AMD HIP: accept gfx1101 alongside gfx1100closed · @jhohertz · 0 评论 · 2026-10-01
- #262 HIP: gfx1201 (Radeon AI PRO R9700) support, faster packed byte intrinsics, and a gfx1201 hipBLASLt tableclosed · @ttio2tech · 0 评论 · 2026-10-01
- #234 Bench: RTX 3090 / dual RTX 3090 on EPYC 7453closed · @mad9home · 0 评论 · 2026-10-01
- #233 docs: add community benchmark guide and RTX 5090 resultsclosed · @hagope · 0 评论 · 2026-10-01
- #231 fix(serve): report a tool call cut off by the end of the output as unfinishedclosed · @alphastorm · 0 评论 · 2026-10-01
- #186 decode: hyper-connection read in 2 kernels, stream-split (-1.5 ms per verify window)closed · @q8atnight · 0 评论 · 2026-10-01
- #242 IQ kernels: decode each weight part once for every column and entry, bitwise identicalclosed · @gputier · 0 评论 · 2026-10-01
- #241 Faster grouped and per-hit Q2_0 expert kernels, bitwise identicalclosed · @gputier · 0 评论 · 2026-10-01
- #258 fused_gr: TILE=1280 kernel specialization for sm_75 (Turing) down-projectionclosed · @hireymage · 0 评论 · 2026-10-01
- #264 tests: make iq_parity reproducible from a fresh checkoutclosed · @j-luwierski · 0 评论 · 2026-10-01
- #240 gemm: name the cuBLAS/CUDA call that failed at init, and flag resource shortagesclosed · @lukmanfauzie · 0 评论 · 2026-10-01
- #235 verify: STRATA_LOGPOS hook for per-position log-probabilities (items 12-13 of #84)closed · @enkynakamura · 0 评论 · 2026-10-01
- #129 core: add optional shared expert arena backingclosed · @rhgo1749 · 0 评论 · 2026-09-30
- #96 Add docker buildclosed · @djmaze · 0 评论 · 2026-09-30
- #269 Split prefill: run the whole prompt on the main GPU (--prefill-main)closed draft · @mijkathegreat · 0 评论 · 2026-09-30
- #265 nvme kv cache v2 - the delta tier + v4 snapshots (replaces #52)closed · @maedoc · 0 评论 · 2026-09-30
- #190 Spill evicted conversation snapshots to a bounded optional disk cacheclosed draft · @jeremiahritchey · 0 评论 · 2026-09-30
- #237 Use a second GPU's VRAM as an opt-in expert store for the prompt pathclosed · @lukmanfauzie · 0 评论 · 2026-09-30
- #247 HIP: Windows and gfx1201 (RX 9070)closed · @jagsan-cyber · 0 评论 · 2026-09-30
- #260 RDNA3 WMMA kernels for the gfx1100 prefill (dense GEMMs and prompt attention)closed · @StevenChenSE · 0 评论 · 2026-09-30
- #228 QSA select on tensor cores (rocWMMA, opt-in, gfx1100) + AMD RX 7900 XTX rates row + verify profiler columns fixclosed · @xyzzing · 0 评论 · 2026-09-30
- #168 hip: support Q2_K embeddings and type-aware PLE table rowsclosed · @samuelishida · 0 评论 · 2026-09-30
- #239 verify: stage the mode-2 PCIe share on the copy stream inside the captured windowclosed · @hireymage · 0 评论 · 2026-09-30
- #238 Let the native router and the expert geometry take the count from the fileclosed · @lukmanfauzie · 0 评论 · 2026-09-30
- #255 Q8 0 supportclosed · @gopinath87607 · 0 评论 · 2026-09-30
- #245 Doubled throughputclosed · @eddoursul · 0 评论 · 2026-09-30
- #227 AVX1: the Q2_0 expert kernels for CPUs that have AVX but no AVX2closed · @demetree · 0 评论 · 2026-09-30
- #223 Multi gpu: Layer split: pinned arena, lent prompt buffers, tested planner (2 x 2080 Ti decode +34%, 29K prompts +61%)closed · @giostrives · 0 评论 · 2026-09-30
- #216 Multi-GPU: the session carve and the per-stage prompt loansclosed · @gopinath87607 · 0 评论 · 2026-09-30
- #124 Pascal (sm_60): bring the engine up on compute capability 6.0closed · @ruibeikaa · 0 评论 · 2026-09-30
- #208 server: optional idle unload, /unload and /load, a free-VRAM guard (share the GPU with other programs)closed · @bytethecookie · 0 评论 · 2026-09-30
- #181 Windows: the engine, vision encoder and MCP servers end with the server (closing the window no longer orphans them)closed · @Apposite245 · 0 评论 · 2026-09-30
- #52 nvme kv cacheclosed · @maedoc · 0 评论 · 2026-09-30
- #148 tests: bitwise check of native_mmvq's multi_exact contract, with a negative controlclosed · @enkynakamura · 0 评论 · 2026-09-30
- #167 generate: zero session state for native packs before the prompt pathclosed · @samuelishida · 0 评论 · 2026-09-30
- #163 feat: show prefill speed alongside decode speed in dashboardclosed · @hendrikp · 0 评论 · 2026-09-30
- #158 WSL2: survive the driver's ~1 GiB pinned host budget with 3 GPUsclosed · @MrRoza · 0 评论 · 2026-09-30
- #225 perf(prefill): honor explicit routed-only staging for large chunksclosed · @rluisr · 0 评论 · 2026-09-30
- #202 ple: --ple-io ram locks the n-gram table in RAM (no SSD read on the prompt/token path)closed · @q8atnight · 0 评论 · 2026-09-30
- #197 sampler: top_k by lexicographic threshold, split over the whole GPUclosed · @gputier · 0 评论 · 2026-09-30
- #187 decode: QSA block scores read each key block once per window (bit-identical)closed · @q8atnight · 0 评论 · 2026-09-30
- #188 prefill: software-pipelined loads in the column-split GDN recurrence (bit-identical)closed · @q8atnight · 0 评论 · 2026-09-30
- #154 Correctness fixes from #149 (s_gemv barrier, bf16 NaN, QSA page mask, verify-window bounds, MSVC, test build)closed · @gputier · 0 评论 · 2026-09-30
- #139 perf(prefill): add shape-selected pre-Ampere BF16 SGEMM with current-main A/Bclosed · @rluisr · 0 评论 · 2026-09-30
- #203 serve: prompt checkpoints copied out asynchronously (pinned staging on the prompt stream + a thread)closed · @q8atnight · 0 评论 · 2026-09-30
- #166 build: replace CUDART_INF_F with INFINITY for HIP-only toolchains (fixes #157)closed · @samuelishida · 0 评论 · 2026-09-30
- #110 Multi-GPU: a second GPU as an expert tier for decode and prompts (--peer-device), + multi-conversation cacheclosed · @q8atnight · 0 评论 · 2026-09-30
- #192 HIP: support RDNA3 gfx1101/gfx1102 (RX 7600/7700/7800), not only gfx1100closed · @nexus2905 · 0 评论 · 2026-09-30
- #177 Web UI: prefill speeds, request timings, a context gauge, chat compactionclosed draft · @orangeswim · 0 评论 · 2026-09-30
- #176 HIP backend on gfx1200 (RDNA4, RX 9060 XT): the gfx1100 backend runs unchanged with three deltas (Changes made by GLM5.3-Flash from freebuff)closed · @Efeisot · 0 评论 · 2026-09-30
- #170 hip: gfx1100 hipBLASLt tuning table for ROCm 7.2.1 (100202)closed · @samuelishida · 0 评论 · 2026-09-30
- #169 generate/verify: make diagnostic dumps work under --spec for native packsclosed · @samuelishida · 0 评论 · 2026-09-30
- #151 Add secondary GPU expert store (--vram-experts) and low-RAM multi-GPU setupclosed · @lukmanfauzie · 0 评论 · 2026-09-30
- #140 Add OrcaRouter Qwen3.8 Flash-Next Q4_K_S supportclosed · @Suoriks · 0 评论 · 2026-09-30
- #130 setup: Volta (V100, sm_70) and Turing as an experimental build pathclosed · @DingoOz · 0 评论 · 2026-09-30
- #105 The Monitor's GPU readings on ROCm: an amdgpu sysfs backend, and the disk rate without psutilclosed · @xyzzing · 0 评论 · 2026-09-30
- #80 Add --tiered-experts: run with less RAM than the experts need (32 GB works)closed · @andrewcoul · 0 评论 · 2026-09-30
- #209 Feature/support CUDA <12.3 compatibility closed · @giostrives · 0 评论 · 2026-09-30
- #153 Expert cache keeps ~4 GiB more on WDDM (+10% decode on a 32 GB card); Ctrl+C and client hang-ups on Windowsclosed · @borexola · 0 评论 · 2026-09-30
- #205 fix(engine): clear stale error/cancellation string between requests and before prompt prefillclosed · @praveshkhatana · 0 评论 · 2026-09-30
- #194 serve: clear the error string at the start of every request (fixes #183)closed · @demon851113 · 0 评论 · 2026-09-30
- #175 Preserve independent conversation caches across interleaved requestsclosed · @b7216309-jpg · 0 评论 · 2026-09-30
- #185 sampler: the sampled path over 64 vocabulary chunks per token (identical picks, ~12% faster sampled decode)closed · @q8atnight · 0 评论 · 2026-09-30
- #87 Turing port: run on RTX 20 (sm_75) and newerclosed · @hireymage · 0 评论 · 2026-09-30
- #180 Sampler: add llama.cpp-compatible DRY penaltiesclosed draft · @bearice · 0 评论 · 2026-09-30
- #102 feat: optional RAM and VRAM guard for the server it launchesclosed draft · @midhatn · 0 评论 · 2026-09-29
- #100 fix: honor native per-layer layouts in file-backed expert loadingclosed · @midhatn · 0 评论 · 2026-09-29
- #41 serve: preserve primary conversation state across auxiliary callsclosed draft · @midhatn · 0 评论 · 2026-09-29
- #160 HIP follow-up: native-pack fixes, type-aware PLE rows, and clean test provisioningclosed · @samuelishida · 0 评论 · 2026-09-29
- #101 build: allow sm_70 (Volta) builds — the only sm_80+ dependency is unreferencedclosed · @eightman999 · 0 评论 · 2026-09-29
- #121 Add gfx1100 HIP support with tuned prefill and bounded expert uploadsclosed · @2jztricks · 0 评论 · 2026-09-29
- #120 `--kv k8v4`: hybrid KV cache — INT8 K + Hadamard-rotated Q4_0 V (816 B/cell)closed · @orangeswim · 0 评论 · 2026-09-29
- #111 Add secondary GPU expert store (--vram-experts) for low-RAM workstatins but with multi-GPUclosed · @lukmanfauzie · 0 评论 · 2026-09-29
- #109 decode: batch the verify window's per-token kernels (bit-identical, +~10% decode)closed · @q8atnight · 0 评论 · 2026-09-29
- #126 Windows: the expert load runs 24x slower from Task Scheduler (0.05 vs 1.42 GiB/s); a hint names the causeclosed · @Zauberio · 0 评论 · 2026-09-29
- #122 serve: add llama.cpp-compatible discovery endpointsclosed · @hendrikp · 0 评论 · 2026-09-29
- #118 serve: routing traces from the live serverclosed · @maedoc · 0 评论 · 2026-09-29
- #113 serve: keep image parts in tool/assistant messagesclosed · @Yunado · 0 评论 · 2026-09-29
- #104 The Monitor's live tok/s is a rate over the last seconds, not the mean since the first tokenclosed · @xyzzing · 0 评论 · 2026-09-29
- #92 setup: name a missing or short shard with its numbers; verify reads back the manifest sha256closed · @Avicennasis · 0 评论 · 2026-09-29
- #91 GGUF reader: refuse a duplicate tensor name at open, naming the tensor and the fileclosed · @Avicennasis · 0 评论 · 2026-09-29
- #89 loader: bulk-read the expert arena and the MTP drafter (MSVC splits ifstream reads into 4095-byte calls)closed · @dannychirkov · 0 评论 · 2026-09-29
- #82 Web app works behind path-prefixed reverse proxiesclosed · @mikicvi · 0 评论 · 2026-09-29
- #81 Accept sm_80+ GPUs at runtime, matching the build guard and the READMEclosed · @mikicvi · 0 评论 · 2026-09-29
- #114 serve: report prefix reuse as usage.prompt_tokens_details.cached_tokensclosed · @Yunado · 0 评论 · 2026-09-29
- #83 serve: llama-server-style timings on responses (prefill/decode rates for llama-swap & friends)closed · @mikicvi · 0 评论 · 2026-09-29
- #88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)closed · @hireymage · 0 评论 · 2026-09-29
- #108 prefill: bit-identical kernel speed-ups (batched indexer append, parallel GDN conv, column-split GDN recurrence, one-launch embedding gather)closed · @q8atnight · 0 评论 · 2026-09-29
- #79 linux: make DirectFile reads actually asynchronousclosed · @andrewcoul · 0 评论 · 2026-09-29
- #94 Add experimental gfx1100 HIP backend for RX 7900 XTXclosed · @2jztricks · 0 评论 · 2026-09-29
- #16 Add optional multi-GPU expert caches and compact transfersclosed draft · @Daerdaal · 0 评论 · 2026-09-28
- #65 cache: preserve root system prompt checkpoint across turns and evictionclosed · @code-martin · 0 评论 · 2026-09-28
- #69 telemetry: per-request decode expert cache hit rate in DONE and web monitorclosed · @code-martin · 0 评论 · 2026-09-28
- #67 Support OrcaRouter IQ3_XXS with explicit BF16 compatibility packingclosed · @acrogenesis · 0 评论 · 2026-09-28
- #62 serve: the conversation cache keeps its shared prefix; the rest rotates LRUclosed · @j-luwierski · 0 评论 · 2026-09-28
- #63 Fix Windows PermissionError in get_llama_cpp()closed · @alfirus · 0 评论 · 2026-09-28
- #50 tools: add mmproj quantization utilityclosed · @code-martin · 0 评论 · 2026-09-28
- #39 telemetry: decode-only cache hit rate, prompt checkpoint retention, and dashboardclosed · @code-martin · 0 评论 · 2026-09-28
- #38 moe: continuous adaptive expert cache eviction with exponential decayclosed · @code-martin · 0 评论 · 2026-09-28
- #37 core: elastic expert cache via CUDA VMM and on-demand GPU vision lifecycleclosed · @code-martin · 0 评论 · 2026-09-28
- #61 text sampler: an empty penalty window must not touch the unsized shared bitmapclosed · @j-luwierski · 0 评论 · 2026-09-28
- #59 Fix/53 sampler penalties. sampler: penalties apply once, as in llama.cpp's default chain (#53); guard the unsized penalty bitmapclosed · @j-luwierski · 0 评论 · 2026-09-28
- #54 Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)closed · @pjgmobile · 0 评论 · 2026-09-28
- #19 serve: per-request temperature/top_p/top_k/min_p/penalties samplingclosed · @j-luwierski · 0 评论 · 2026-09-27
- #15 generate: residual control-vector projection from GGUFsclosed · @j-luwierski · 0 评论 · 2026-09-27
- #24 serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejectingclosed · @pjgmobile · 0 评论 · 2026-09-27
- #22 tools: add real-time web dashboard for hardware and engine monitoringclosed · @code-martin · 0 评论 · 2026-09-27
- #21 kv: add Q4_0 KV cache mode with FWHT-256 Hadamard rotationclosed · @code-martin · 0 评论 · 2026-09-27
- #18 serve: treat unset/-1 max_tokens as remaining context, not 1024closed · @coolio986 · 0 评论 · 2026-09-27
- #10 serve: read short prompt parts through the decode windows (~1 s faster first token)closed · @Mirtraxxx · 0 评论 · 2026-09-26
- #8 Conversation cache: keep the chat between requests, read only what is newclosed · @Mirtraxxx · 0 评论 · 2026-09-26
- #7 serve: fix zero-token replies from engine-queue race between requestsclosed · @coolio986 · 0 评论 · 2026-09-26
- #9 CPU pool: sleep between requests instead of spinning (fixes #4)closed · @Mirtraxxx · 0 评论 · 2026-09-26
- #14 CPU pool: fix dangling else in Linux physical_cores() (fixes #13, root cause of #11)closed · @Vistawizard · 0 评论 · 2026-09-26
- #3 V100 (sm_70) support for the Qwen3.8-Flash-Next MoE engineclosed · @noorazman · 0 评论 · 2026-09-25