Pull requests / #1420

#1420 Add optional GLM-5.3-Flash support from Project Maya

open · draft · @apr3ndi5 · 0 Kommentare · Auf GitHub

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux

Beschreibung

## Summary

Add GLM-5.3-Flash as an optional model family alongside Qwen, using Project Maya at commit [444030a3f2afc01515d4d067d663377329efb5b5](https://github.com/mw00/project-maya/commit/444030a3f2afc01515d4d067d663377329efb5b5) as the source reference. This branch is based on Strata main at d5ea7133741e67743c0e886bb426c0ce8d69cf6c (0.1.40.3).

This is a draft pending CUDA compilation, device parity checks, and full-model validation on Linux/NVIDIA and experimental native Windows/NVIDIA hardware.

## What changed

- Add a separate, opt-in GLM engine path (`STRATA_ENABLE_GLM=ON`, `--glm-pack`) with the Maya model architecture, tokenizer, CUDA kernels, VRAM/RAM/SSD expert tiers, two-GPU layer splitting, and MTP support. Qwen remains the default.
- Extend the existing installer for native Linux x86-64 and experimental native Windows x64 with NVIDIA: Maya-S v2 IQ2_XXS selection, pinned SHA-256 verification, packing, source builds, configuration, and launch script generation. An unattended model download requires explicit `--download-model` authorization.
- Integrate GLM chat formatting, image processing and GPU memory lending into the existing Strata server and web app, retaining the OpenAI and Anthropic API formats. Add expert-tier statistics to the Monitor.
- Require a nonempty API key before binding the server to an external address, including configuration-based startup.
- Add offline installer, pack, tokenizer, oracle and server checks, synthetic CUDA parity fixtures, and documentation in `docs/GLM.md`. Preserve the applicable MIT and OFL notices.

## Initial-port validation (f2f680d; before this update)

Completed locally on Windows on 2026-10-07:

- 479 Python test cases completed: 472 passed and 7 skipped, covering mock server/API behavior, installer, pack, tokenizer, oracle, configuration, detokenization and Qwen regressions.
- 3 C++ cache/native-helper tests passed.
- The CPU image encoder, including the GLM projector, built with MSVC 19.51.
- Changed Python files compiled, JavaScript syntax checks passed, and the Git whitespace check passed.
- Enabling the GLM engine build on Windows was rejected as expected.

Still required on a supported Linux/NVIDIA machine:

- CUDA compilation/device linking and the prepared device parity tests.
- Full-model text and image inference, expert-tier behavior, GPU memory lending/reclaim, and one/two-GPU MTP validation.

No model weights were downloaded and no production model engine or server was started for these checks. Mock server tests were run.

## Extra Notes

The GLM engine and kernels are ported from the pinned Maya reference with adaptations for Strata's build and internal APIs; this is not a byte-for-byte copy of the Maya repository. In particular, `STRATA_GLM_SLOW=1` uses diagnostic CPU experts; the older Maya slow GPU expert pool has not been ported. The default GLM fast path includes the GPU/RAM/SSD tiers.

No speed or quality measurements from Maya are claimed as measurements of this integration. GLM targets native Linux/NVIDIA and experimental native Windows x64/NVIDIA. WSL2, AMD and Intel remain unsupported for this model family.

### Update to Project Maya v1.2.0

Port the changes from Project Maya commit `444030a` through `5932f601373f53fc021f75dc55159a722c772571`.

- Improve GLM prefill with a deeper SSD landing buffer, next-layer expert read-ahead, and GPU-sized prompt chunks.
- Add experimental native Windows/NVIDIA support through Strata's existing installer, including asynchronous expert reads and commit-aware pinned-memory sizing.
- Preserve Strata's Qwen default, API formats, dashboard, download verification, and external-host API-key requirement.

No performance measurements from Maya are presented as measurements of this Strata integration. Full-model inference remains unverified in this checkout.

#### Fresh validation for this update

Run locally on Windows on 2026-10-07 with Python 3.10.5 and MSVC 19.51 (Visual Studio 2026):

- 590 Python test cases completed: **583 passed, 7 skipped**. This includes 331 passed/5 skipped mock server/API/frontend/security/lifecycle/monitor/VRAM/detokenization cases and 252 passed/2 skipped installer/GLM pack/tokenizer/oracle/configuration/responses/Qwen-regression cases.
- The 16 GLM installer cases also passed after final adjustments. Coverage includes native Linux and Windows, compatible Visual Studio discovery, WSL/AMD/Intel/ARM rejection, `.bat` launchers, configuration, external-host API keys, and explicit consent for model/vision downloads. Reruns are not counted twice in the total above.
- **4 C++ tests passed**: conversation cache, conversation memory, native split, and the new prefill-memory helper (13 checks for sizing, retries and fallback).
- The synthetic fixture generated and packed successfully on CPU: 184 tensors, 15 Q4_0 expert tensors and 5 disk-backed MoE layers. This checks the fixture and pack, not GPU inference.
- The CPU image encoder, including its GLM projector, built successfully. The existing llama.cpp MSBuild duplicate-source warning remains.
- Changed Python files compiled; JavaScript syntax and Git whitespace checks passed.
- CMake rejected GLM with Intel/SYCL and 32-bit Windows. Native Windows x64 GLM configuration reached CUDA discovery and failed with `No CUDA toolset found`.

Prepared, but **not executed**: `glm_prefill_parity`, using five synthetic Q4_0 prompt chunks, SSD reads, 12/64 landing slots and read-ahead off/on. It compares 8,192 continuation-logit values with token-at-a-time execution at maximum absolute error below `2e-3`. Minimal GLM Q4_0 fast-kernel/MMQ support is included for this fixture; its CUDA compilation and numerical behavior are not verified here.

CUDA compilation/device linking, device parity, actual CUDA pinned-allocation failure/cleanup, and full-model text/image/MTP inference remain pending on compatible NVIDIA/CUDA systems. This machine has no CUDA toolkit; Visual Studio 2026 does not satisfy the recommended CUDA 12.8/Visual Studio 2022 combination.

No model weights were downloaded, installed models/configurations/environments were retained, and no production model engine or server was started. The PR remains a draft.

Final update commit: `cd2f4a9f540dfa94c06a0823bfacc8d2e7cad355`.

Mehr auf der Site

Links zu Install, Modellen, Releases.