贡献 / #1420

#1420 Add optional GLM-5.3-Flash support from Project Maya

open · draft · @apr3ndi5 · 0 评论 · 去 GitHub 看

Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux

说明

## Summary



Add GLM-5.3-Flash as an optional model family alongside Qwen, using Project Maya at commit [444030a3f2afc01515d4d067d663377329efb5b5](https://github.com/mw00/project-maya/commit/444030a3f2afc01515d4d067d663377329efb5b5) as the source reference. This branch is based on Strata main at d5ea7133741e67743c0e886bb426c0ce8d69cf6c (0.1.40.3).



This is a draft pending CUDA compilation, device parity checks, and full-model validation on Linux/NVIDIA and experimental native Windows/NVIDIA hardware.



## What changed



- Add a separate, opt-in GLM engine path (`STRATA_ENABLE_GLM=ON`, `--glm-pack`) with the Maya model architecture, tokenizer, CUDA kernels, VRAM/RAM/SSD expert tiers, two-GPU layer splitting, and MTP support. Qwen remains the default.

- Extend the existing installer for native Linux x86-64 and experimental native Windows x64 with NVIDIA: Maya-S v2 IQ2_XXS selection, pinned SHA-256 verification, packing, source builds, configuration, and launch script generation. An unattended model download requires explicit `--download-model` authorization.

- Integrate GLM chat formatting, image processing and GPU memory lending into the existing Strata server and web app, retaining the OpenAI and Anthropic API formats. Add expert-tier statistics to the Monitor.

- Require a nonempty API key before binding the server to an external address, including configuration-based startup.

- Add offline installer, pack, tokenizer, oracle and server checks, synthetic CUDA parity fixtures, and documentation in `docs/GLM.md`. Preserve the applicable MIT and OFL notices.



## Initial-port validation (f2f680d; before this update)



Completed locally on Windows on 2026-10-07:



- 479 Python test cases completed: 472 passed and 7 skipped, covering mock server/API behavior, installer, pack, tokenizer, oracle, configuration, detokenization and Qwen regressions.

- 3 C++ cache/native-helper tests passed.

- The CPU image encoder, including the GLM projector, built with MSVC 19.51.

- Changed Python files compiled, JavaScript syntax checks passed, and the Git whitespace check passed.

- Enabling the GLM engine build on Windows was rejected as expected.



Still required on a supported Linux/NVIDIA machine:



- CUDA compilation/device linking and the prepared device parity tests.

- Full-model text and image inference, expert-tier behavior, GPU memory lending/reclaim, and one/two-GPU MTP validation.



No model weights were downloaded and no production model engine or server was started for these checks. Mock server tests were run.



## Extra Notes



The GLM engine and kernels are ported from the pinned Maya reference with adaptations for Strata's build and internal APIs; this is not a byte-for-byte copy of the Maya repository. In particular, `STRATA_GLM_SLOW=1` uses diagnostic CPU experts; the older Maya slow GPU expert pool has not been ported. The default GLM fast path includes the GPU/RAM/SSD tiers.



No speed or quality measurements from Maya are claimed as measurements of this integration. GLM targets native Linux/NVIDIA and experimental native Windows x64/NVIDIA. WSL2, AMD and Intel remain unsupported for this model family.



### Update to Project Maya v1.2.0



Port the changes from Project Maya commit `444030a` through `5932f601373f53fc021f75dc55159a722c772571`.



- Improve GLM prefill with a deeper SSD landing buffer, next-layer expert read-ahead, and GPU-sized prompt chunks.

- Add experimental native Windows/NVIDIA support through Strata's existing installer, including asynchronous expert reads and commit-aware pinned-memory sizing.

- Preserve Strata's Qwen default, API formats, dashboard, download verification, and external-host API-key requirement.



No performance measurements from Maya are presented as measurements of this Strata integration. Full-model inference remains unverified in this checkout.



#### Fresh validation for this update



Run locally on Windows on 2026-10-07 with Python 3.10.5 and MSVC 19.51 (Visual Studio 2026):



- 590 Python test cases completed: **583 passed, 7 skipped**. This includes 331 passed/5 skipped mock server/API/frontend/security/lifecycle/monitor/VRAM/detokenization cases and 252 passed/2 skipped installer/GLM pack/tokenizer/oracle/configuration/responses/Qwen-regression cases.

- The 16 GLM installer cases also passed after final adjustments. Coverage includes native Linux and Windows, compatible Visual Studio discovery, WSL/AMD/Intel/ARM rejection, `.bat` launchers, configuration, external-host API keys, and explicit consent for model/vision downloads. Reruns are not counted twice in the total above.

- **4 C++ tests passed**: conversation cache, conversation memory, native split, and the new prefill-memory helper (13 checks for sizing, retries and fallback).

- The synthetic fixture generated and packed successfully on CPU: 184 tensors, 15 Q4_0 expert tensors and 5 disk-backed MoE layers. This checks the fixture and pack, not GPU inference.

- The CPU image encoder, including its GLM projector, built successfully. The existing llama.cpp MSBuild duplicate-source warning remains.

- Changed Python files compiled; JavaScript syntax and Git whitespace checks passed.

- CMake rejected GLM with Intel/SYCL and 32-bit Windows. Native Windows x64 GLM configuration reached CUDA discovery and failed with `No CUDA toolset found`.



Prepared, but **not executed**: `glm_prefill_parity`, using five synthetic Q4_0 prompt chunks, SSD reads, 12/64 landing slots and read-ahead off/on. It compares 8,192 continuation-logit values with token-at-a-time execution at maximum absolute error below `2e-3`. Minimal GLM Q4_0 fast-kernel/MMQ support is included for this fixture; its CUDA compilation and numerical behavior are not verified here.



CUDA compilation/device linking, device parity, actual CUDA pinned-allocation failure/cleanup, and full-model text/image/MTP inference remain pending on compatible NVIDIA/CUDA systems. This machine has no CUDA toolkit; Visual Studio 2026 does not satisfy the recommended CUDA 12.8/Visual Studio 2022 combination.



No model weights were downloaded, installed models/configurations/environments were retained, and no production model engine or server was started. The PR remains a draft.



Final update commit: `cd2f4a9f540dfa94c06a0823bfacc8d2e7cad355`.


## Strata 0.1.41 conflict resolution — 2026-10-08

Merged official main `fb58e0dbc8399662c0e47c76578c6e878b14f6cf` into this branch without rewriting its published history. Resolved setup CUDA-toolkit-root/GLM flags, baseline-ISA native expert predicates/GLM buffer capacities, and Qwen piece-cache/GLM ignore-merges conflicts. Qwen remains the default.

Fresh checks: 38 focused installer, tokenizer, GLM API and request-hardening cases passed; the native-split CPU target rebuilt and passed on MSVC 19.51. These are fresh merge checks, separate from the previous port results. CUDA/device/full-model inference is still pending; no model weights were downloaded or production engine/server started.

Conflict-resolution commit: `ff9f6b8832b4d12aba40f8553204d120b7a3cc05`. The newer Maya updates are in draft PR #1476.

本站相关内容

相关页面的快捷入口。