Release notes / v0.1.25

Strata v0.1.25

Downloads

Release notes

Faster prompts, AMD Radeon support (experimental), and a smaller KV cache option.

Faster prompts (same output): reading a prompt got 7% faster. RTX 5070, Q2_0:

  • 32K tokens: 1,901 → 2,030 tokens/s.
  • 128K tokens: 1,955 → 2,100 tokens/s.

Two changes:

  • The per-layer expert grouping tables no longer wait behind the expert stream on the copy engine.
  • The hyper-connection kernels are fused: no FP32 copy of the normalized rows, and each write is fused with the next norm.

The model's output is bit-identical to 0.1.24.

AMD Radeon RX 7900 XT / XTX on Linux (experimental): ./setup.sh --backend hip (chosen by itself when the PC has no NVIDIA card Strata can use).

  • It finds the card through the kernel's amdgpu driver.
  • It installs ROCm into Strata's own Python environment (no sudo) and compiles the engine on your PC.
  • It uses GEMM kernels tuned for the 7900 XTX.
  • Coder IQ1_M on an RX 7900 XTX: prompts ~1,200-1,330 tokens/s at 8-32K, answers 50-68 tokens/s.
  • For now: one GPU, no images. See docs/AMD_HIP.md.
  • The backend is PR #121, thanks @2jztricks. It succeeds #94.

`--kv k8v4`: a smaller KV cache (optional): INT8 keys and 4-bit rotated values, 23% less VRAM for the context. The freed VRAM holds more experts, so long conversations answer faster. On an RTX 3090 at 198K: 85 → 99 tokens/s, with the same needle results. Prompts are 2-5% slower than int8. It can't be combined with KV streaming, which setup turns on from 64K. PR #120, thanks @orangeswim.

Checked before the release:

  • Byte-identical to 0.1.24 on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, Coder), including the prompt path's internal state at 4K and 20K.
  • Builds and runs the same on Linux.
  • k8v4's parity tests pass.
  • AMD: the HIP test suite, a live server test of every endpoint, and needle tests on the RX 7900 XTX.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.25.

The ready-made Strata engine for Windows (RTX 30 / 40 / 50: sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.

Full notes on GitHub

Newer release: v0.1.26Older release: v0.1.24