Notes de version / v0.1.28

Strata v0.1.28

Téléchargements

Notes de version

Fixes: VRAM left free again after 0.1.27's larger draft head, a cancelled request no longer kills the next one, and the API key protects every endpoint.

VRAM (#199, #201, #217). The expert cache takes the free VRAM minus a reserve (700 MiB) that keeps the GPU from paging. The draft layer's head is allocated after the cache, and 0.1.27's larger head (with Chinese, Japanese and Korean) came out of that reserve: on a 16 GB card with images, 156 MiB were left, and requests could stall. The cache now reserves the draft head as well, so the reserve is kept. Auto-sized caches hold a few percent fewer experts than in 0.1.27.

  • New: START-HERE.bat --setup --draft-vocab en keeps the English/code draft subset from before 0.1.27: ~110 MiB less

VRAM and English answers 1-2% faster, but Chinese, Japanese and Korean answers get almost no drafts.

  • When no VRAM is left for the expert cache at all, the engine now says what to lower instead of a bare "verify:"

error (#174).

Cancelled requests (#183). A request cancelled while its prompt was being read (a client that disconnects, a timeout) left an error behind, and the next long prompt then ended the engine, which the server had to reload. Fixed by @praveshkhatana (#205); thanks to @demon851113 for the root cause and the validation (#194).

Security.

  • /status showed the end of the answer being written without asking for the API key. It needs the key now, and it

no longer keeps the answer's end after the request (#212).

  • The key is compared in constant time, and an explicitly empty key (--api-key "" or an empty STRATA_API_KEY)

stops the server instead of silently switching authentication off (#213).

  • --host and --api-key given when starting an installed model were ignored. They are now saved for that model

(#179).

Tool calls (#210). A tool call whose arguments contain </parameter> or </tool_call> (for example a file that documents the chat format) was cut off at that text. A call now ends only at its own closing tags.

Setup and server:

  • When the engine exits, the server shows its last log line and exit code. Before, it always blamed the RAM (#215).
  • RTX 50 cards: a compiled engine needs CUDA 13.0. An engine built with 12.8 crashed in the prompt path on Linux

(#220).

  • Windows: llama.cpp's web UI is not unpacked. Its long paths broke setup in deep folders like

Downloads\Strata-main\Strata-main (#206).

  • Swift 1.5 Q2_0 is hidden for now: its files split one layer across the two shards, which the pack tool cannot

prepare yet (#171). Swift IQ2_XS has a similar speed.

  • --data-dir also moves the files of the data folder used before and of a nested Strata-data folder (#198).
  • Model files copied in by hand are recognized when they are whole, instead of being downloaded again (#173).
  • The expert cache on WDDM keeps ~4 GiB more on a 32 GB card, Ctrl+C stops the server on Windows, and a client that

hangs up no longer prints a stack trace (#153, @borexola).

  • CUDA older than 12.3 builds again (#209, @giostrives).

Checked before the release:

  • Byte-identical to 0.1.27 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),

including the prompt path's internal state.

  • The no-repeat and sanity checks with the default settings, and needle tests at 8K, 16K and 32K.
  • A cancelled long prompt followed by the same prompt and a short one: the engine stays up.
  • The server's tests (68, including the new ones for /status, the key and tool-call values).
  • Builds and runs the same on Linux.
  • On this PC's Q2_0 setup, 379 MiB of VRAM are left free with everything loaded (0.1.27: 210 MiB, below the stall

line).

  • AMD: not tested again for this release (our AMD test PC was not reachable). The changes there are small shared

C++ edits.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux: ./setup.sh). Setup installs engine 0.1.28.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.

Notes complètes sur GitHub

Version plus récente: v0.1.29Version plus ancienne: v0.1.27