Release notes / v0.1.39
Strata v0.1.39
Downloads
- NVIDIA · CUDA 12 (167 MiB)
- AMD · HIP (571 MiB)
- NVIDIA · CUDA 13 (118 MiB)
Release notes
Faster decode and long prompts, several requests at once, older and other hardware as experimental opt-ins, the OpenAI Responses API for Codex, a fix for slower prompts with a RAM budget, and a batch of fixes from your reports.
Faster (the same answers where nothing says otherwise):
- Decode (#646): the verify pass runs with fewer launches and host round trips (sub-warp expert packing, staged
inputs, the PLE and MTP steps batched). Measured here on one RTX 5070 against 0.1.38, 10 interleaved pairs, medians: Q2_0 +6% (story 68.4 -> 72.7 tok/s, code 75.8 -> 80.0), IQ3_XXS +6% / +2.5% (46.1 -> 48.8, 50.7 -> 51.9); after 4K and 32K prompts +5% to +7%. The output is byte-identical to 0.1.38's (10/10 on all four quants with a fixed cache). The zero-doorbell verify graph itself runs only when every expert of a layer is in VRAM. A 12 GB card never gets there, so the gain on big cards was not measured here.
- Long prompts (#583, on by default): the streamed expert ring is sized in bytes for the pack, and
--prefill auto
picks the largest chunk that keeps it full. Short prompts, and prompts that fit 0.1.38's chunk, keep 0.1.38's ring and their exact bits (4K: identical). The gain depends on the free VRAM. On an RTX 5070 with 32K prompts, IQ3_XXS is +18.5% with a 1,500-slot expert cache and unchanged with setup's default config. A long prompt's bits change against 0.1.38, because the experts go through a different mix of cached and streamed groups. Its quality stays in the same band, teacher-forced over the next 2,001 tokens against the FP16 prompt path (IQ3_XXS): 8K prompt KL 0.042 (0.1.38: 0.054), top-1 93.5% (92.2%); 32K prompt KL 0.020 (0.019), top-1 95.3% (95.0%). It doesn't help every pack: the Coder at a 32K context reads a 30K prompt about 8% slower (with a 64K context it was about 6% faster). STRATA_RING_BYTES=0 restores 0.1.38's ring.
- Very long contexts: the attention's top-k past the register kernel's reach (262K-524K cells) runs with a
histogram per warp: a 243K-token prompt +26% on an RTX 3060 (#603).
- More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
--remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out. A layer split of 3+ GPUs overlaps its stages better (#598, 4x RTX 3060: +10 to 28% on prompts). STRATA_PREFILL_HELP=1 lets a split's idle card stream a share of a one-chunk prompt's experts (#663). It is opt-in because it rounds differently. With one or two GPUs the pipeline runs as before.
- Also: the Linux expert arena on transparent huge pages (#650;
STRATA_NO_LARGEPAGES=1keeps 4 KB pages), an
opt-in AVX2 codebook gather for the IQ CPU kernels (STRATA_IQ256_GATHER=1, #622), and thread affinity on PCs with more than 64 CPUs (#626).
Several requests at once (#465, opt-in): "parallel": N in strata-<model>.json (or setup --parallel N) decodes up to N conversations together in the engine. More requests wait for a free slot, and a long prompt gives way to a short one at a chunk boundary. A request left alone goes back to the one-at-a-time path, and /metrics shows each slot. Every slot's greedy answer is the same as when it runs alone. On a 12 GB card it cuts waiting time and costs speed. RTX 5070, Q2_0, four requests at once: the last one starts after 1.8 s instead of 11.2 s, but together they decode 11% slower (63.1 against 70.7 tok/s). A request alone loses 11% with 2 slots and 22% with 4, because each slot's session takes 0.56 GiB from the expert cache. So setup recommends it only where the experts mostly fit in VRAM. A 4-GPU layer split (4x 16 GB, IQ3_S, PR #559) served 8 requests at 360 tok/s in total against 120 for one. docs/BATCHING.md has the numbers and the options.
Older and other hardware, experimental: three opt-in paths that community members wrote and measured on their own machines. We have none of this hardware. Each is compile-checked and unit-tested here, and the ready-made engines and their output are unchanged (checked byte-identical).
- Older NVIDIA cards (Pascal, Volta: P40, P100, GTX 10, V100, Titan V; #395 #600 #540 #655 #627): CUDA 13 cannot
compile for them, so setup keeps a second engine built with CUDA 12.9 (strata-windows-x64-cuda12.zip on Windows, compiled on Linux). It is used only when you choose such a card: a PC with only Pascal / Volta cards, a card named with --gpu N / --gpus, or --cuda 12. --cuda 12 is also the way to run with an NVIDIA driver older than 580 (528+ on Windows, 525+ on Linux). The choice is kept per model. The Volta prompt attention and the BF16 path through FP16 / fp32 are compiled into this engine only. Untested here: every Pascal and Volta card, the old drivers, and an RTX 50 card in that engine (it gets a warning; keep it on its own model). docs/OLDER_GPUS.md has the reporters' numbers (V100: prompts 1,123-1,251 tok/s, UD-IQ4_XS). RTX 20 owners can try the same FP16 tensor-core path for the prompt's BF16 products with STRATA_BF16_TC=1 (+15-18% on an RTX 2080 Ti, #655). Its sums are not bitwise cuBLAS's. Older AMD cards are built by hand: gfx906 (Instinct MI50 / MI60, Radeon VII; -DSTRATA_HIP_GFX906=ON, #638 #677: 2x MI50, the Coder, 50 tok/s at 4K) and gfx1012 (RX 5500 XT, #442). The RX 6700 XT (gfx1031, #524) goes through setup.
- Intel Arc (#423): maxfridbe's SYCL port of the engine, built from source on Linux with Intel oneAPI:
./setup.sh --backend sycl. Reported on the Arc Pro B70 / B50 and the B580 (Coder IQ1_M 70-78 tok/s on a B70) on earlier versions. Here it compiles (oneAPI 2026.1) and its kernel tests run on a CPU device. Untested: the 0.1.39 port on an Arc, Windows (no build path yet; setup points at Linux), WSL2, the A-series and integrated Arc GPUs, and images. There is no ready-made Intel engine. docs/INTEL_ARC.md.
- Older CPUs without AVX2 (#394 #595): AVX-only (Sandy / Ivy Bridge, Xeon E5 v1/v2, Bulldozer) and
SSE4.2-only CPUs (Nehalem / Westmere). Run setup as usual: it warns and compiles the engine for that CPU (STRATA_ISA_FLOOR=avx or none, 10-20 minutes once) instead of stopping. Only the i-quant models run there, and the CPU's share is slow (forced on our Ryzen: 11-17 tok/s with the AVX build, 3.5-3.8 with SSE4.2, against 26). Untested here on a real old CPU; contributors ran earlier versions of it on Xeon E5-2680 / E5-2687W / X5690. docs/INSTALL.md#older-cpus-experimental.
Prompts with a RAM budget are fast again (#577): 0.1.38 decided too early whether to read the experts past the file cache, and it compared against the size of every model file. On a 96 GB PC with UD-Q4_K_XL and a 72 GiB budget it chose wrong, so every refill after a prompt read the drive (prompts 15-40% slower than 0.1.34). The choice now counts only the expert bytes outside the RAM copy, and is made again once that copy is built. The tokens are the same either way.
The OpenAI Responses API (#451): POST /v1/responses, so Codex CLI works with Strata (tested with Codex 0.160.0, a tool loop included; later turns reused ~96% of the prompt from the cache). It is stateless, as Codex uses it, and covers function tools, tool results, reasoning effort, JSON schemas and the streaming events. Not supported: previous_response_id (nothing is stored), hosted tools and reasoning summaries. The Codex config.toml is in docs/DETAILS.md.
Security: SECURITY.md says how to report a problem privately (GitHub's private vulnerability reporting) and what the server exposes: 127.0.0.1 by default, the API key, the Host and Origin checks of 0.1.38, CORS, and the opt-in MCP tools and request monitor.
Fixes from your reports:
- A reply stuck on one token is ended (#606): a reply that re