Notas de versão / v0.1.26

Strata v0.1.26

Downloads

Notas da versão

Prompts up to 32K read 4-10% faster: the draft layer's prompt pass runs in batches.

Before answering, Strata runs its draft layer over the prompt, so the drafter (MTP speculative decoding) knows the conversation. That pass used to run 6 positions at a time: about 0.9 s of a 32K prompt. Now it runs in batches through the prompt path's kernels: 0.9 s → 0.07 s.

  • 32K prompts are 5-6% faster, 12K prompts 4-10%.
  • Drafts are as good as before. Over six varied 12K prompts, the drafts accepted: Q2_0 0.700 → 0.699, the Coder 0.716 → 0.718. Answer speed is unchanged.
  • The model's own numbers are untouched: this changes only the draft layer's memory of the prompt. The generated text can still differ from 0.1.25 in rare cases, when a different draft makes the verifier batch differently (the same kind of last-bit difference as a different speculation depth).
  • With KV streaming (setup turns it on from 64K), the draft layer keeps its cache in a ring and uses the old pass for now.
  • STRATA_MTP_BATCH=0 switches back.

Low-RAM mode: a big GPU makes up for little RAM. Normally every expert of the model is copied into RAM (23-50 GB), even when the GPU could hold most of them. On a 32 GB PC even the Coder did not start (an RTX 5090 user). Now, when the model's experts would not fit the RAM beside the system, setup maps them from one file in the model's folder instead (--mmap-experts):

  • The OS file cache holds what the GPU does not, and it can give that memory back.
  • On the Coder, the engine's committed memory drops from 36 GB to about 13 GB, with the same answers.
  • With a 32 GB card, the Coder's experts all sit on the GPU, and Q2_0's mostly do.
  • With a small GPU, most experts come from the SSD: much slower, and setup says so.
  • --low-ram on|off overrides the choice. See "Low-RAM mode" in docs/DETAILS.md.

Setup: a Linux PC with both NVIDIA and AMD cards now asks which to use. Before, it took the NVIDIA cards without mentioning the Radeon.

Checked before the release:

  • With the batched pass switched off, byte-identical to 0.1.25 on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, Coder), including the prompt path's internal state.
  • With it on: the no-repeat and sanity checks on every quant, and needle tests at 8K, 16K and 32K.
  • Builds and runs the same on Linux.
  • On the RX 7900 XTX (AMD): the HIP test suite (30/30), the batched pass (prompts +13-19%), and a live server test.
  • The low-RAM mode: a full install with it, the same answers as the normal mode, and a live server test.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.26.

The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.

Completo no GitHub

Versão mais nova: v0.1.27Versão anterior: v0.1.25