标签 / Security
Security
Auto topic: Security
Issues
- #1495 [Feature]: experimental per-stage KV-grow for a two-GPU layer split at native 262K
- #1478 gfx1201 (R9700): the gfx1151 prompt switches in arch_defaults.cpp are exact and +4-8% prompt; suggest a gfx12 entry (decode half is slower)
- #1470 Pascal (sm_61): decode halves since 0.1.40.2. 2e4ddf6 drops `__restrict__` and peels the MMVQ loop on every CUDA arch
- #1276 UPDATE.bat cannot update existing checkouts after the main history rewrite
- #1244 Docker v0.1.40: persisted /data/config edits ignored on restart; stale /opt/strata config used (worked in v0.1.39)
- #1212 Review: shared MTP head fix — sibling Q4 state omission and test configuration gaps
- #1177 Feature request: add timestamps to engine log lines
- #1054 [SYCL / Intel Arc] 2x Arc Pro B70 --layer-split auto deadlocks at startup after the host-mirror fill
- #1005 CRLF line-break tokens (`\r\n`, ids 317/845/8301) get far too little probability; everything else matches llama.cpp
- #916 Unable to launch Strata on RTX 2060 6GB VRAM + 32 GB RAM
- #735 can't load vision encoder
- #661 Proposal: save and restore a conversation to disk (/slots/0?action=save|restore)
- #654 HIP (gfx1200) RuntimeError Exception Code: 0xC0000005
- #629 Re-running setup with a different --context silently discards mcp_servers, mcp and sampling from the run config
- #613 AMD 9070 XT on Windows causes VIDEO_ENGINE_TIMEOUT_DETECTED
- #544 Please run several security audits
- #537 Literal </think> quoted in reasoning yields empty stop or leaked reasoning (v0.1.37)
Pull requests
- #1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMA
- #1493 Experimental cooperative memory relief and resumable text requests
- #1480 serve: --conversation-cache-disk-only - the conversation cache on disk, with no RAM budget (depends on #1271, #1269)
- #1476 Update GLM port to Project Maya v1.0.4
- #1449 Add Metal backend support for Strata on macOS
- #1446 Responses: replay fixes, experimental persistence and optional summaries
- #1420 Add optional GLM-5.3-Flash support from Project Maya
- #1405 tools: add configurable launches and reuse local model downloads
- #1394 serve: evict parked conversations to pass the physical-RAM gate instead of dropping the snapshot
- #1382 docs: index entry for the TITAN RTX community report
- #1343 Add support for CYBER-FROST-3.8 (Blackfrost): read qwen4exp.nextn_predict_layers
- #1326 serve: support IPv6 bind addresses
- #1323 prefill: batched unbuffered stager reads; close the experts.bin view once reads are unbuffered (port of #833)
- #1314 serve: keep unfinished tool calls visible as content
- #1297 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #1269 serve: restore a session file without holding its K/V in RAM
- #1265 Add the engine updater: verify, stage, all-or-nothing swap, rollback
- #1260 tests: order MMQ parity transfers on the compute stream
- #1251 --peer-device prompt share without P2P: host route, per-token sums, FP16 transfers (2.1-2.5x prompts on a no-P2P pair)
- #1231 kv-grow: a run that cannot lend cache slots maps the whole window up front instead of writing past its first 16K cells
- #1202 dashboard: Serve the web page under an optional dashboard path
- #1152 Fix Windows AMD telemetry and add optional Arabic UI
- #1144 serve: add opt-in instruction skills to Chat
- #1136 Add opt-in durable chat history, legacy import and browser compaction
- #1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
- #1126 Harden server, installer and engine against high-risk bugs
- #1119 Add local PDF, Word and Excel attachments with screenshot previews
- #1032 serve: add opt-in bounded Responses history and continuation
- #992 security-hardening
- #969 serve: Q4 at 297.3 tok/s non-MTP (N=8); MTP +30.4% decode (N=2)
- #898 perf: Q8 resident adaptation (+24.5–28.2% rotation gain on RTX PRO 6000)
- #886 serve: skip completed empty assistant turns in chat history
- #861 serve: "strata_checkpoint": false lets a one-shot request skip its conversation checkpoint (#830)
- #858 gfx906 compat: cudaFuncSetAttribute as a template function, the shared-memory carveout name (#646's fused_gr did not build)
- #810 serve: accept json_schema roots that are unions of object schemas
- #797 sycl: enable Arc A770 inference and document measured performance
- #783 perf(cuda): fused decode/verify/MTP kernels and graph launch reductions on 0.1.39
- #759 Add native Responses API support for Codex clients
- #753 serve: isolate concurrent request metrics (depends on #559)
- #752 cache: preserve checkpoints when split allocation fails (depends on #559)
- #751 Persist conversation KV state in a bounded disk LRU cache
- #744 HIP: on-device autotuning for gfx1100 decode (kernel shapes, draft and CPU settings)
- #727 Add opt-in durable chat history, legacy import and browser compaction
- #724 Add local PDF, Word and Excel attachments with screenshot previews
- #719 Add sandboxed HTML previews and scrollable chat code blocks
- #701 serve: a malformed "tools" is a 400 naming the field (#592)
- #683 setup: a re-run of setup carries over the hand-edited config blocks (#629)
- #680 Add persistent chat history, local model selection, and a Windows desktop client
- #668 Session files: save and restore a conversation to disk (POST /slots/0?action=save|restore)
- #645 Update the engine from the web app, cautiously
- #630 serve: add opt-in stateless Responses API
- #594 serve: read the request body an answer left unread, so the close is not a reset
- #572 serve: return unfinished tool call text as content
- #571 Feat/responses api 451
- #569 serve, setup: an empty API key is refused from the config file and from setup's --api-key (#213)
- #484 Monitor: show the conversation cache at work (reuse, hits on switches, slots, cache RAM)
- #482 Networked Pool support
- #332 serve: add a standalone API request monitor
- #321 serve: add CORS preflight (OPTIONS), reverse-proxy header support, and X-Accel-Buffering for SSE
- #190 Spill evicted conversation snapshots to a bounded optional disk cache
- #189 Preserve alternating conversations with a shared snapshot core and bounded RAM cache
- #122 serve: add llama.cpp-compatible discovery endpoints
- #88 Pascal port: lower the CUDA floor to sm_61 (GTX 10 series)
- #54 Support pruned expert variants (GSQ-RCO-Coder: 256 of 512 experts)
- #24 serve: optional --fit-max-tokens clamps the output cap to remaining context instead of rejecting
- #15 generate: residual control-vector projection from GGUFs
- #12 Fork: uncensored model choices (OrcaRouter, mradermacher, RVN) in the…