Notes de version / v0.1.40
Strata v0.1.40
Téléchargements
- NVIDIA · CUDA 12 (181 MiB)
- AMD · HIP (598 MiB)
- NVIDIA · CUDA 13 (126 MiB)
Notes de version
Strix Halo (gfx1151) support, crash, NaN and security fixes, better multi-GPU, batching and resident modes, Q4_0 / Q4_1 and more PLE table formats, AMD fixes, a much faster Linux cold start, and a long list of opt-ins.
Strix Halo / Ryzen AI Max (gfx1151, experimental): the HIP engine now runs on the Radeon 8060S / 8050S, and setup recognizes the chip (PCI 1002:1586), sizes memory for it as one shared pool, recommends UD-IQ4_XS, and on Linux picks HIP by itself. The engine has RDNA3.5 matrix-core kernels for the prompt, a hipBLASLt table for ROCm 7.14.1, free-memory accounting for unified-memory chips (#409, a DGX Spark gets it too) and 18 speed switches that turn on by themselves on gfx1151 (STRATA_GFX1151_DEFAULTS=0 turns them off). All of them gave byte-identical output here. The Windows HIP zip now includes gfx1151 code. It is experimental, and we have not run it on Windows Strix Halo hardware. docs/STRIX_HALO.md has the toolchain (no root needed), a manual build and the opt-in prompt switches. The 8040S mapping in setup is a guess. Measured on a Ryzen AI Max+ 395 with 128 GB, UD-IQ4_XS, medians of interleaved runs:
| Context | Prompt tok/s | Output tok/s | |---|---|---| | 8K | 1,293 | 53.8 | | 64K | 1,370 | 46.6 | | 128K | 1,320 | 51.4 |
That is +6% prompt and +4% output over our earlier engine branch. UD-Q4_K_XL, IQ3_S and IQ3_XXS are in the doc. The table matches the final defaults, which keep the shared-expert stream fork on for gfx1151 (it is the 18th automatic switch): UD-IQ4_XS at 8K gave 54.07 output against 53.94 in the table. Thanks to StevenChenSE for the RDNA3 matrix-core GEMM kernels (#313).
Crash and NaN fixes:
- Non-finite values (#838, #550, #871, #606, #879): the fused SwiGLU q8_1 quantizers keep their scale finite. The residency-table upload waits for its own copy before the verifier reads it. The all-resident verify graph runs only while every expert is in VRAM, and a stale plan fails the window instead of printing
!!!!. A Stop sent during a prompt read reaches the engine at once. We could not reproduce #879 here, so it stays open for a retry on 0.1.40. - `--batch` on a split (#776, #845): a fully resident split stage no longer waits for doorbells that are never rung.
- A dead engine (#748): a fatal verify timeout is caught before batch routing, and the next request is refused.
- Parking a conversation (#752): a host out-of-memory skips the park instead of ending the server.
- HIP: the shared-expert stream fork is off by default on AMD, which brings decode back to 0.1.38 speed (#826, #816). The doorbell kernels fence after the ring store (#697), which may explain the batched-prompt stalls in #579 and #613.
- Windows: the RAM budget respects the commit limit (#749, #730), a large-page arena is no longer VirtualLocked (#779), and Smart App Control blocks get a clear message (#735).
Security:
- #553: network image paths are refused, URLs are capped at 32 MiB, and a page from another origin cannot name a local file as a picture. We added the same check to
/v1/responses, which the PR did not cover. - We did not merge #434 (a Manager page that any website can drive, plus a proxy that adds the API key) or #696 (the API key becomes a shell on a LAN server). The replies say what each needs.
More than one GPU, batching and resident modes:
- Resident RAM mode on a layer split (#848, thanks to Francesco Albano): about 70 tok/s against 32-64 on 32 GB machines. It also fixes a 0.1.39 bug where the adaptive swaps copied back from the wrong card's cache. Setup keeps the resident variant on a split only with an engine that has it, so setup now needs engine 0.1.40.
- A helper GPU's experts are not in the PCIe share (#854): 40 -> 66 tok/s on a helper rig, the author's number.
- Four GPUs (#887): the layer split scores every four-way placement instead of guessing. This changes the automatic placement on 4 GPUs. 2 and 3 GPUs keep 0.1.39's, and
STRATA_SPLIT_COVER_Basks for the new one at any count. - Stage weights (#880):
--layer-split autocan trim the later stages' weights (4 GPUs, prompt 470 -> 1,930 tok/s, the author's number).--vram-reserve-later-mibsets a separate reserve for the later cards. - Batching:
--batch-mtpkeeps MTP drafting in concurrent slots (#846, +31 to +39% with 2-4 clients, the author's number).serveruns without--mtp(#708). - Fusions ported from Eddoursul's fork: the draft round runs the draft layer's front once, a one-token window commits itself, the draft kernels use the main layer's forms, the final mixer is one fused read, and the argmax runs over many blocks.
--host-core first|lastand the--spec-follow/--window-hashestest tools come from the same fork. The #783 kernel work (QSA early exit, GDN decode kernels, vec4 router, batched KV append, multi-token GR) is merged with a kill switch for each. Its long-context QSA early exit gives +2.7% decode at 120K (6 pairs).STRATA_GR_DOWN_MAX4stays opt-in, because it gave no gain on the 5070. - Decode against 0.1.39 on NVIDIA (RTX 5070, 12 pairs, 200-token medians): Q2_0 code -0.3%, story +3.4%; IQ3_XXS code +2.9%, story +2.3%. Prompts are within 2%.
New formats:
- Q4_0 and Q4_1 experts (#599) on the GPU kernels, and a Q4_0 PLE table.
- PLE tables in Q8_0, Q5_1 and BF16 (#651, #586, #865): one table of PLE formats now covers these and Q5_0 and FP8, with bounds checks on the file. Against the BF16 table, Q8_0 is off by 0.53%, FP8 by 2.64% and IQ4_NL by 7.60%. The effect on answers is small (KL about 0.012 for any of them), so setup does not offer a bigger table.
--ple-io ramfaults the table in on 16 threads (8 GB cold in WSL: 15 s -> 6 s). - Unsloth UD-Q6_K_XL: the Q8_0 PLE table is read, and the Q6_K expert kernels are an opt-in build (
-DSTRATA_Q6K_EXPERTS=ON, #860). - IQ3_XXS and IQ4_XS PLE keys stay native, so those packs start again (#381).
- `--kv k8v4` streams with `--kv-resident` (#711, #705), and setup offers it.
AMD:
- RDNA2 (#835, opt-in):
STRATA_HIP_PROMPT_F16=1runs the prompt's 16-bit GEMMs in FP16, and prompts read about 2x faster on an RX 6900 XT. It rounds differently, so it is off for now. The engine prints a tip on gfx103x. - hipBLASLt tables for gfx1100 (#766, #755, #565), RDNA3 / RDNA4 sibling arch names (#695), gfx1034 (#778) and a gfx906 build fix (#808).
- Windows HIP (#873, #654): one
hipFree(0)warm-up before the first memory query. - Opt-in:
STRATA_HIP_ADAPT_KERNEL_COPY=1does the adaptive swaps' copies with a kernel instead of SDMA (#884, for the 2x gfx1030 hang).STRATA_DENSE_MMQ=1(#820) gave no gain on gfx1151 and stays opt-in. - Older cards: the Q4_0 expert kernel overflowed on sm_60 (signed overflow). It is fixed, found on a new P100 test box.
- Intel Arc (#784, #866, #868, #870): the SYCL port compiles against 0.1.39's sources again, a ring wait over 2 ms is no longer reported as finished, and the Arc Pro B60 ids are known.
Linux cold start: weights, head, embedding, vision encoder and the RAM copy are now read ahead of use (#699). On an RTX 5090 with 32 GB, Q2_0 at 262K, ready in 70 s instead of about 920 s. STRATA_READ_AHEAD=0 turns it off. Linux also gets the unbuffered file tier (#773, O_DIRECT, same rule as Windows). On a 16 GB laptop with 30 GB RAM, decode went 4.4 -> 7.9 tok/s and startup 36 -> about 20 s (author's numbers). With transparent huge pages on always the arena no longer asks for more (#771, STRATA_NO_ARENA_THP=1).
Defaults that changed (the rest of the default output is unchanged):
- Empty assistant turns (#886): they are not rendered into the next prompt, so a client's history can't teach the model to answer with nothing.
STRATA_KEEP_EMPTY_TURNS=1restores the old rendering. - `tool_choice` (#790):
noneoffers no tools on both routes.requiredand a named function work. - `json_object` mode (#762): JSON is taken o
Version plus récente: v0.1.40.1Version plus ancienne: v0.1.39