リリースノート

Synced 2026-10-08 · 30 versions · 最新 v0.1.41

Strata v0.1.41 最新

Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group…

リリースノート · GitHub で全文

Strata v0.1.40.4

Hotfix for 0.1.40.3: decode on Pascal cards (Tesla P40, GTX 1080, GTX 1070) is back to the 0.1.40 speed. Default answers are unchanged. Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with…

リリースノート · GitHub で全文

Strata v0.1.40.3

Hotfix for 0.1.40.2: Intel Arc install and A-series fixes, Windows AMD speed and runtime fixes, Docker config links, and a few smaller fixes. Default answers are unchanged. Update with UPDATE.bat (Linux: ./update.sh);…

リリースノート · GitHub で全文

Strata v0.1.40.2

Faster and fixed across the board: official Intel Arc support (Arc Pro B70 and A750 tested), bug fixes and speed-ups merged from the open issues and PRs, and big gains on multi-GPU, low-RAM Linux and long-document use.…

リリースノート · GitHub で全文

Strata v0.1.40.1

Hotfix for 0.1.40: a tool call that is only quoted (in the thinking or in a code block) no longer runs, and requests waiting during an engine restart no longer hang. Python server fixes only. The engine is the same as…

リリースノート · GitHub で全文

Strata v0.1.40

Strix Halo (gfx1151) support, crash, NaN and security fixes, better multi-GPU, batching and resident modes, Q4_0 / Q4_1 and more PLE table formats, AMD fixes, a much faster Linux cold start, and a long list of opt-ins.…

リリースノート · GitHub で全文

Strata v0.1.39

Faster decode and long prompts, several requests at once, older and other hardware as experimental opt-ins, the OpenAI Responses API for Codex, a fix for slower prompts with a RAM budget, and a batch of fixes from your…

リリースノート · GitHub で全文

Strata v0.1.38

Faster prompts, --kv q4_0 prompts on tensor cores, a 6 GB card that starts, and a large batch of community PRs. Security: the server without an API key (DNS rebinding, cross-site requests). Without api_key, the server…

リリースノート · GitHub で全文

Strata v0.1.37

A frozen engine no longer hangs the server, Windows AMD counts the desktop's VRAM, and a batch of setup fixes from your reports. The server notices a silent engine (481): if the engine prints nothing during a request…

リリースノート · GitHub で全文

Strata v0.1.36

Faster: Q2_0 reads prompts 16-22% faster and writes up to 18% faster at long contexts. Plus one-click updates (UPDATE.bat), an opt-in memory of which experts you use, and clearer logs. Speed (136), measured on an RTX…

リリースノート · GitHub で全文

Strata v0.1.35

Windows AMD no longer crashes on the first prompt, the low-RAM mode works on 32 GB Windows PCs, and a batch of fixes from your reports. AMD on Windows (468, 461): the engine now uses the HIP runtime it ships with.…

リリースノート · GitHub で全文

Strata v0.1.34

AMD cards on Windows (new), an MCP server so your AI assistant can set Strata up, a shorter README - and a cancelled request now frees the engine within a second. AMD on Windows (new, 247 325): an AMD card on Windows is…

リリースノート · GitHub で全文

Strata v0.1.33

The image encoder runs on every CPU again, and setup recommends instead of forcing: a longer context, more GPUs or a bigger RAM budget you choose is kept, with a note about the risk. Images (411, 412): 0.1.32's…

リリースノート · GitHub で全文

Strata v0.1.32

Multi-GPU short prompts fixed, AMD decode +12%, Unsloth's Q4 model in setup with 2.8x faster prompts, and 30+ community changes - with the same answers. Multi-GPU: short prompts faster again (340, riverhh76). Since…

リリースノート · GitHub で全文

Strata v0.1.31

Unsloth's Q4 model runs on a 64 GB PC (experimental), AMD decode up to 15% faster, models load twice as fast on Windows, and a batch of fixes. Experimental: Unsloth UD-Q4_K_XL on a normal PC. Strata can now run…

リリースノート · GitHub で全文

Strata v0.1.30

Short prompts up to 28% faster (the same answers), multi-GPU and AMD RDNA4 improvements, a server that can give the GPU back, and several opt-in features. Faster short prompts. Prompts of 1K to 4K tokens read all…

リリースノート · GitHub で全文

Strata v0.1.29

Sampled answers up to 40% faster (the same text), faster prompt kernels, correctness fixes, and several contributions. Sampled answers (197, @gputier). With a temperature above 0 (most chat apps), each token's top-k…

リリースノート · GitHub で全文

Strata v0.1.28

Fixes: VRAM left free again after 0.1.27's larger draft head, a cancelled request no longer kills the next one, and the API key protects every endpoint. VRAM (199, 201, 217). The expert cache takes the free VRAM minus a…

リリースノート · GitHub で全文

Strata v0.1.27

Answers in Chinese, Japanese or Korean are written faster, RTX 20 cards are supported, and five fixes. Faster CJK answers (137). The draft layer (MTP speculative decoding) can only propose tokens from its subset of the…

リリースノート · GitHub で全文

Strata v0.1.26

Prompts up to 32K read 4-10% faster: the draft layer's prompt pass runs in batches. Before answering, Strata runs its draft layer over the prompt, so the drafter (MTP speculative decoding) knows the conversation. That…

リリースノート · GitHub で全文

Strata v0.1.25

Faster prompts, AMD Radeon support (experimental), and a smaller KV cache option. Faster prompts (same output): reading a prompt got 7% faster. RTX 5070, Q2_0: - 32K tokens: 1,901 → 2,030 tokens/s. - 128K tokens: 1,955…

リリースノート · GitHub で全文

Strata v0.1.24

Long prompts read faster: the sparse attention's selection runs on tensor cores. For every token of a prompt, Strata picks which earlier parts of the context its attention reads (the QSA selection). At long context that…

リリースノート · GitHub で全文

Strata v0.1.23

Faster decoding, llama.cpp-compatible server endpoints, and fixes for images, 8 GB cards and multi-GPU. Faster answers: the verify window's per-token kernels (router, norms, RoPE, the shared-expert cast, the combine)…

リリースノート · GitHub で全文

Strata v0.1.22

Faster prompts, and several GPUs without typing any flags. Prompts are read faster (RTX 5070, Q2_0): a 32K-token prompt 1,213 → 1,646 tokens/s, a 128K prompt with KV streaming 1,144 → 1,596 tokens/s (time to first token…

リリースノート · GitHub で全文

Strata v0.1.21

One model across two or three NVIDIA cards (experimental). - Layer split across GPUs: - Setup: START-HERE.bat --setup --gpus 0,2 (Linux: ./setup.sh --setup --gpus 0,2), or "gpu": [0, 2] in a run config. - Each card runs…

リリースノート · GitHub で全文

Strata v0.1.20

Faster new chats for agent clients, a hit-rate column in the Monitor, a PCIe-aware default, and community fixes. - New chats reuse the system prompt (62 by @j-luwierski, 65 by @code-martin): - A prompt read from the…

リリースノート · GitHub で全文

Strata v0.1.19

Penalties that really apply, tuning for your PC, tools from MCP servers, and a fix for starting with a small page file. - Fixed: repetition, presence and frequency penalties were mostly not applied (53). - What went…

リリースノート · GitHub で全文

Strata v0.1.18

Clearer engine updates, a SETUP.bat shortcut, and a sampler guard. - Updating right after a release no longer keeps the old engine without saying why (58, thanks @XeonG). New files on main could reach you a few minutes…

リリースノート · GitHub で全文

Strata v0.1.17

Claude Code works with Strata, sampling matches llama.cpp, and you can choose the GPU. - Claude Code can use Strata (55, 56, thanks @Nicolas0315): the server accepts /v1/messages?beta=true, and a system message in the…

リリースノート · GitHub で全文

Strata v0.1.16

New: the Coder model. No more re-downloads after an update. Strata is now open source (MIT). - The GSQ-RCO Coder ([ISTA-DASLab](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF)), thanks to…

リリースノート · GitHub で全文