Notas de versão / v0.1.40

Strata v0.1.40

Downloads

Notas da versão

Strix Halo (gfx1151) support, crash, NaN and security fixes, better multi-GPU, batching and resident modes, Q4_0 / Q4_1 and more PLE table formats, AMD fixes, a much faster Linux cold start, and a long list of opt-ins.

Strix Halo / Ryzen AI Max (gfx1151, experimental): the HIP engine now runs on the Radeon 8060S / 8050S, and setup recognizes the chip (PCI 1002:1586), sizes memory for it as one shared pool, recommends UD-IQ4_XS, and on Linux picks HIP by itself. The engine has RDNA3.5 matrix-core kernels for the prompt, a hipBLASLt table for ROCm 7.14.1, free-memory accounting for unified-memory chips (#409, a DGX Spark gets it too) and 18 speed switches that turn on by themselves on gfx1151 (STRATA_GFX1151_DEFAULTS=0 turns them off). All of them gave byte-identical output here. The Windows HIP zip now includes gfx1151 code. It is experimental, and we have not run it on Windows Strix Halo hardware. docs/STRIX_HALO.md has the toolchain (no root needed), a manual build and the opt-in prompt switches. The 8040S mapping in setup is a guess. Measured on a Ryzen AI Max+ 395 with 128 GB, UD-IQ4_XS, medians of interleaved runs:

| Context | Prompt tok/s | Output tok/s | |---|---|---| | 8K | 1,293 | 53.8 | | 64K | 1,370 | 46.6 | | 128K | 1,320 | 51.4 |

That is +6% prompt and +4% output over our earlier engine branch. UD-Q4_K_XL, IQ3_S and IQ3_XXS are in the doc. The table matches the final defaults, which keep the shared-expert stream fork on for gfx1151 (it is the 18th automatic switch): UD-IQ4_XS at 8K gave 54.07 output against 53.94 in the table. Thanks to StevenChenSE for the RDNA3 matrix-core GEMM kernels (#313).

Crash and NaN fixes:

  • Non-finite values (#838, #550, #871, #606, #879): the fused SwiGLU q8_1 quantizers keep their scale finite. The residency-table upload waits for its own copy before the verifier reads it. The all-resident verify graph runs only while every expert is in VRAM, and a stale plan fails the window instead of printing !!!!. A Stop sent during a prompt read reaches the engine at once. We could not reproduce #879 here, so it stays open for a retry on 0.1.40.
  • `--batch` on a split (#776, #845): a fully resident split stage no longer waits for doorbells that are never rung.
  • A dead engine (#748): a fatal verify timeout is caught before batch routing, and the next request is refused.
  • Parking a conversation (#752): a host out-of-memory skips the park instead of ending the server.
  • HIP: the shared-expert stream fork is off by default on AMD, which brings decode back to 0.1.38 speed (#826, #816). The doorbell kernels fence after the ring store (#697), which may explain the batched-prompt stalls in #579 and #613.
  • Windows: the RAM budget respects the commit limit (#749, #730), a large-page arena is no longer VirtualLocked (#779), and Smart App Control blocks get a clear message (#735).

Security:

  • #553: network image paths are refused, URLs are capped at 32 MiB, and a page from another origin cannot name a local file as a picture. We added the same check to /v1/responses, which the PR did not cover.
  • We did not merge #434 (a Manager page that any website can drive, plus a proxy that adds the API key) or #696 (the API key becomes a shell on a LAN server). The replies say what each needs.

More than one GPU, batching and resident modes:

  • Resident RAM mode on a layer split (#848, thanks to Francesco Albano): about 70 tok/s against 32-64 on 32 GB machines. It also fixes a 0.1.39 bug where the adaptive swaps copied back from the wrong card's cache. Setup keeps the resident variant on a split only with an engine that has it, so setup now needs engine 0.1.40.
  • A helper GPU's experts are not in the PCIe share (#854): 40 -> 66 tok/s on a helper rig, the author's number.
  • Four GPUs (#887): the layer split scores every four-way placement instead of guessing. This changes the automatic placement on 4 GPUs. 2 and 3 GPUs keep 0.1.39's, and STRATA_SPLIT_COVER_B asks for the new one at any count.
  • Stage weights (#880): --layer-split auto can trim the later stages' weights (4 GPUs, prompt 470 -> 1,930 tok/s, the author's number). --vram-reserve-later-mib sets a separate reserve for the later cards.
  • Batching: --batch-mtp keeps MTP drafting in concurrent slots (#846, +31 to +39% with 2-4 clients, the author's number). serve runs without --mtp (#708).
  • Fusions ported from Eddoursul's fork: the draft round runs the draft layer's front once, a one-token window commits itself, the draft kernels use the main layer's forms, the final mixer is one fused read, and the argmax runs over many blocks. --host-core first|last and the --spec-follow / --window-hashes test tools come from the same fork. The #783 kernel work (QSA early exit, GDN decode kernels, vec4 router, batched KV append, multi-token GR) is merged with a kill switch for each. Its long-context QSA early exit gives +2.7% decode at 120K (6 pairs). STRATA_GR_DOWN_MAX4 stays opt-in, because it gave no gain on the 5070.
  • Decode against 0.1.39 on NVIDIA (RTX 5070, 12 pairs, 200-token medians): Q2_0 code -0.3%, story +3.4%; IQ3_XXS code +2.9%, story +2.3%. Prompts are within 2%.

New formats:

  • Q4_0 and Q4_1 experts (#599) on the GPU kernels, and a Q4_0 PLE table.
  • PLE tables in Q8_0, Q5_1 and BF16 (#651, #586, #865): one table of PLE formats now covers these and Q5_0 and FP8, with bounds checks on the file. Against the BF16 table, Q8_0 is off by 0.53%, FP8 by 2.64% and IQ4_NL by 7.60%. The effect on answers is small (KL about 0.012 for any of them), so setup does not offer a bigger table. --ple-io ram faults the table in on 16 threads (8 GB cold in WSL: 15 s -> 6 s).
  • Unsloth UD-Q6_K_XL: the Q8_0 PLE table is read, and the Q6_K expert kernels are an opt-in build (-DSTRATA_Q6K_EXPERTS=ON, #860).
  • IQ3_XXS and IQ4_XS PLE keys stay native, so those packs start again (#381).
  • `--kv k8v4` streams with `--kv-resident` (#711, #705), and setup offers it.

AMD:

  • RDNA2 (#835, opt-in): STRATA_HIP_PROMPT_F16=1 runs the prompt's 16-bit GEMMs in FP16, and prompts read about 2x faster on an RX 6900 XT. It rounds differently, so it is off for now. The engine prints a tip on gfx103x.
  • hipBLASLt tables for gfx1100 (#766, #755, #565), RDNA3 / RDNA4 sibling arch names (#695), gfx1034 (#778) and a gfx906 build fix (#808).
  • Windows HIP (#873, #654): one hipFree(0) warm-up before the first memory query.
  • Opt-in: STRATA_HIP_ADAPT_KERNEL_COPY=1 does the adaptive swaps' copies with a kernel instead of SDMA (#884, for the 2x gfx1030 hang). STRATA_DENSE_MMQ=1 (#820) gave no gain on gfx1151 and stays opt-in.
  • Older cards: the Q4_0 expert kernel overflowed on sm_60 (signed overflow). It is fixed, found on a new P100 test box.
  • Intel Arc (#784, #866, #868, #870): the SYCL port compiles against 0.1.39's sources again, a ring wait over 2 ms is no longer reported as finished, and the Arc Pro B60 ids are known.

Linux cold start: weights, head, embedding, vision encoder and the RAM copy are now read ahead of use (#699). On an RTX 5090 with 32 GB, Q2_0 at 262K, ready in 70 s instead of about 920 s. STRATA_READ_AHEAD=0 turns it off. Linux also gets the unbuffered file tier (#773, O_DIRECT, same rule as Windows). On a 16 GB laptop with 30 GB RAM, decode went 4.4 -> 7.9 tok/s and startup 36 -> about 20 s (author's numbers). With transparent huge pages on always the arena no longer asks for more (#771, STRATA_NO_ARENA_THP=1).

Defaults that changed (the rest of the default output is unchanged):

  • Empty assistant turns (#886): they are not rendered into the next prompt, so a client's history can't teach the model to answer with nothing. STRATA_KEEP_EMPTY_TURNS=1 restores the old rendering.
  • `tool_choice` (#790): none offers no tools on both routes. required and a named function work.
  • `json_object` mode (#762): JSON is taken o

Completo no GitHub

Versão mais nova: v0.1.40.1Versão anterior: v0.1.39