Notas de versão / v0.1.41
Strata v0.1.41 Mais recente
Downloads
- NVIDIA · CUDA 12 (189 MiB)
- AMD · HIP (599 MiB)
- NVIDIA · CUDA 13 (131 MiB)
Notas da versão
Faster on more than one GPU, faster short prompts on NVIDIA, about twice as fast prompts on Windows with little RAM, and a long list of fixes. Two defaults change: --batch on a layer split now runs one pipeline group per GPU, and short prompt chunks on one NVIDIA GPU get help from the CPU (this one changes the last bits of an answer; one setting turns it off). Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.41. This release includes the Pascal decode fix from 0.1.40.4.
Speed
Against 0.1.40.4, default settings, interleaved pairs, medians (prompt = time to read it; higher is better in every column):
| Machine, model | Prompt 512 / 1K tokens | Prompt 4K / 20K | Decode | |---|---|---|---| | RTX 3060 12 GB, Linux, IQ3_XXS | +28% / +22% | equal | equal | | Tesla P100 16 GB, Linux, IQ3_XXS | +33% / +25% | equal | equal | | RTX 5070 12 GB, Windows, Q2_0 | equal* | equal | +0.4% (story +4.9%) | | RTX 5070, Windows, IQ3_XXS / IQ3_S with a RAM budget | | +16% / +23-25% | +3% | | 1x Radeon AI PRO R9700, Linux, Q2_0 | equal | equal | equal | | Arc Pro B70 32 GB, Linux, IQ3_S | +13% / +12% | +10% / +10% | equal | | Arc A750 8 GB, Linux, IQ3_XXS | equal | equal | equal |
\* On the RTX 5070 with Q2_0, setup's automatic expert cache already holds about 3,700 experts on the card, so the CPU has almost nothing to take during a short prompt; with a smaller cache (--expert-cache 1500) the same prompts are 15-27% faster. Answers in every row are identical to 0.1.40.4 with STRATA_PREFILL_CPU_SHARE=0 (Q2_0, IQ3_XXS, Coder and IQ3_S on the 5070, including 4K/20K-token prompts; IQ3_XXS on the 3060 and P100; Q2_0 on the R9700; the 10-prompt check on both Intel cards).
- Several clients on several GPUs (default; `--batch` with a layer split).
--batch-groups autois now the default: every GPU stage works on its own group of requests at the same time, instead of one request group walking through the stages while the other cards wait. Total tokens per second at 8 clients: +32% to +68% (2, 3 and 4 GPUs; 4x Radeon AI PRO R9700 in the larger cases), with a better p95 as well. The answers are the same as with one group (the 8-request identity test passes on 2, 3 and 4 GPUs).--batch-groups 1turns it off. The first version of this slowed one long request that ran next to short ones by 39-47%; that is fixed and measured. The server also listens with a backlog of 256 now (STRATA_HTTP_BACKLOG): with 30 to 40 clients opening connections at once, some got "connection reset"; at 40 clients there are now 0 resets (134 tok/s in total on 4x R9700). The engine still serves 8 slots at a time (--batchabove 8 is capped), so with 16 or more clients the rest queue.
| 4x R9700, IQ3_S, --batch 8 | 0.1.40.4 | 0.1.41 | |---|---|---| | 8 clients, total tok/s | 88 | 165 (+87%) | | 8 clients, finish p50 | 17.4 s | 9.3 s | | 40 clients, total tok/s | 88 | 153 (+75%) | | 40 clients, connection resets per run | 14-21 | 0 |
- Short prompts up to a third faster on one NVIDIA GPU (default, changes bits). The CPU, idle while a prompt is read, takes part of the experts the GPU would otherwise stream, for chunks below 1,024 tokens (an agent's tool result, a short follow-up). It helps as much as the GPU has to stream: 0.1.41 against 0.1.40.4 reads a 512-token prompt 28% faster on an RTX 3060 and 33% faster on a Tesla P100 (1,000 tokens: 22% and 25%). On an RTX 5070 with Q2_0, setup's automatic cache already holds most of what a short prompt needs and the gain is 0-2%; with a smaller cache it is 15-27%. It is never slower in our runs. Longer prompts are unchanged. The answers can differ slightly from 0.1.40.4: first-token KL against the share off averages 0.004 (max 0.025).
STRATA_PREFILL_CPU_SHARE=0turns it off and gives the exact 0.1.40.x answers. Only one NVIDIA GPU without--batchslots; layer splits, batch and AMD are as before. Thanks to sergqwer, whose measured on/off check this is. - Windows, model bigger than RAM: prompts about twice as fast (#1323). The prompt reader asks the drive for several experts in one request, and gives up the memory-mapped view of
experts.binonce its reads bypass the cache (while it is mapped, Windows serves those reads one at a time). On the RTX 5070 PC the prompt time fell 45-52%, with the same tokens in 18 of 18 runs; in the release check, IQ3_XXS and IQ3_S with a RAM budget read 4K-20K prompts 16-25% faster on the same PC. On Linux the gain is small (3-4% with unbuffered reads, none through the page cache) and never a loss. Thanks to adonizm, whose #833 this is a port of, and to malloc32 for bringing it back. - A prompt whose experts are already in RAM (#1353). When the GGUF experts read in place are in the page cache (90% or more), the prompt reader uses its few-thread profile instead of the 32-thread SSD one, which had made RAM-resident prompts 4-7x slower for the reporter. On our 4x R9700 box (fast storage, 251 GB RAM) the old path was not slow to begin with, so we measured no change there (identical output); please tell us in #1353 what it does on yours. Where the engine cannot tell (Windows), it behaves as before;
STRATA_STAGER_SSD=0/1forces either. - Intel Arc Pro B70: the expert dequant writes in whole 16-byte runs. 200 -> 423 GB/s of writes in the micro-benchmark, prompt +6%, output bit-identical.
- Intel Arc A750: decode 9.8 -> 12.0 tok/s (+22%). The sampler no longer needs FP64, so the A-series FP64 emulation settings are not needed (docs/INTEL.md).
- Tokenizer cache. A resent agent prompt (89K tokens) is encoded in 0.07 s instead of 0.32 s, the same ids.
- Automatic card order. With an auto layer split, the faster card (more SMs x clock; on Linux AMD from the KFD topology) is put last, the order that measured faster in the reports that asked for it (#1352).
"gpu_order": "as_given"keeps your order. Setup no longer writes--remote-expert-optnext to--pipeline-windows(the first turns the second off). The auto placement also sees each stage's PCIe link (#794, from shy). - Prompt batches fill holes (#793): a request takes a free slot below a busy one of its group first, so a request left alone goes back to the plain single-request path.
Fixes
- A deadlock with `parallel: 2` fixed. Long, long, short requests could hang forever: a request waiting for a slot gave back nothing while it waited, and the slot a paused read held needed those lines.
- A full conversation cache evicts old entries to pass the RAM check instead of dropping the snapshot (#1347; a conversation that holds a pinned shared prefix is never evicted this way). Thanks to alanthinker.
- Hung engines (#1317, #1407). A request whose engine prints nothing for 90 s, uses no CPU, moves no disk bytes and has an idle GPU is ended and restarted (
STRATA_ENGINE_STALL_S, 0 turns it off; needspsutil, which setup installs). An engine that is silent but working is never ended by this. The prompt-stall watchdog now waits, up to 10 times its limit, while the OS is still reading the expert file (STRATA_WATCHDOG_IO_S, 0 turns it off); a real hang stops at the old limit. - `CUDA_LAUNCH_BLOCKING=1` hangs every verify window (#1341, #1383): the engine now says so at start and in the stall report, with the variable named. Thanks to ischencheng.
- `--api-key` takes several keys separated by commas, as llama.cpp does (#1344).
- `setup --calibrate` keeps its earlier measurements when a later engine start fails; it drops that candidate and goes on (#1337). The PCIe sweep no longer ends at 1.0 and visits each value once per round (#1332, #1345, from aly8246).
- A config with a `vision` section but no `--vision` in its args now starts the engine with images (#1322).
- Tool calls written `<function= NAME>` are read as the call NAME (#1430). The vision encoder's start error now names the busy or unavailable GPU, and