贡献 / #1420
#1420 Add optional GLM-5.3-Flash support from Project Maya
open · draft · @apr3ndi5 · 0 评论 · 去 GitHub 看
Setup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsSecurityDocumentationWindowsLinux
说明
## Summary Add GLM-5.3-Flash as an optional model family alongside Qwen, using Project Maya at commit [444030a3f2afc01515d4d067d663377329efb5b5](https://github.com/mw00/project-maya/commit/444030a3f2afc01515d4d067d663377329efb5b5) as the source reference. This branch is based on Strata main at d5ea7133741e67743c0e886bb426c0ce8d69cf6c (0.1.40.3). This is a draft pending CUDA compilation, device parity checks, and full-model validation on Linux/NVIDIA and experimental native Windows/NVIDIA hardware. ## What changed - Add a separate, opt-in GLM engine path (`STRATA_ENABLE_GLM=ON`, `--glm-pack`) with the Maya model architecture, tokenizer, CUDA kernels, VRAM/RAM/SSD expert tiers, two-GPU layer splitting, and MTP support. Qwen remains the default. - Extend the existing installer for native Linux x86-64 and experimental native Windows x64 with NVIDIA: Maya-S v2 IQ2_XXS selection, pinned SHA-256 verification, packing, source builds, configuration, and launch script generation. An unattended model download requires explicit `--download-model` authorization. - Integrate GLM chat formatting, image processing and GPU memory lending into the existing Strata server and web app, retaining the OpenAI and Anthropic API formats. Add expert-tier statistics to the Monitor. - Require a nonempty API key before binding the server to an external address, including configuration-based startup. - Add offline installer, pack, tokenizer, oracle and server checks, synthetic CUDA parity fixtures, and documentation in `docs/GLM.md`. Preserve the applicable MIT and OFL notices. ## Initial-port validation (f2f680d; before this update) Completed locally on Windows on 2026-10-07: - 479 Python test cases completed: 472 passed and 7 skipped, covering mock server/API behavior, installer, pack, tokenizer, oracle, configuration, detokenization and Qwen regressions. - 3 C++ cache/native-helper tests passed. - The CPU image encoder, including the GLM projector, built with MSVC 19.51. - Changed Python files compiled, JavaScript syntax checks passed, and the Git whitespace check passed. - Enabling the GLM engine build on Windows was rejected as expected. Still required on a supported Linux/NVIDIA machine: - CUDA compilation/device linking and the prepared device parity tests. - Full-model text and image inference, expert-tier behavior, GPU memory lending/reclaim, and one/two-GPU MTP validation. No model weights were downloaded and no production model engine or server was started for these checks. Mock server tests were run. ## Extra Notes The GLM engine and kernels are ported from the pinned Maya reference with adaptations for Strata's build and internal APIs; this is not a byte-for-byte copy of the Maya repository. In particular, `STRATA_GLM_SLOW=1` uses diagnostic CPU experts; the older Maya slow GPU expert pool has not been ported. The default GLM fast path includes the GPU/RAM/SSD tiers. No speed or quality measurements from Maya are claimed as measurements of this integration. GLM targets native Linux/NVIDIA and experimental native Windows x64/NVIDIA. WSL2, AMD and Intel remain unsupported for this model family. ### Update to Project Maya v1.2.0 Port the changes from Project Maya commit `444030a` through `5932f601373f53fc021f75dc55159a722c772571`. - Improve GLM prefill with a deeper SSD landing buffer, next-layer expert read-ahead, and GPU-sized prompt chunks. - Add experimental native Windows/NVIDIA support through Strata's existing installer, including asynchronous expert reads and commit-aware pinned-memory sizing. - Preserve Strata's Qwen default, API formats, dashboard, download verification, and external-host API-key requirement. No performance measurements from Maya are presented as measurements of this Strata integration. Full-model inference remains unverified in this checkout. #### Fresh validation for this update Run locally on Windows on 2026-10-07 with Python 3.10.5 and MSVC 19.51 (Visual Studio 2026): - 590 Python test cases completed: **583 passed, 7 skipped**. This includes 331 passed/5 skipped mock server/API/frontend/security/lifecycle/monitor/VRAM/detokenization cases and 252 passed/2 skipped installer/GLM pack/tokenizer/oracle/configuration/responses/Qwen-regression cases. - The 16 GLM installer cases also passed after final adjustments. Coverage includes native Linux and Windows, compatible Visual Studio discovery, WSL/AMD/Intel/ARM rejection, `.bat` launchers, configuration, external-host API keys, and explicit consent for model/vision downloads. Reruns are not counted twice in the total above. - **4 C++ tests passed**: conversation cache, conversation memory, native split, and the new prefill-memory helper (13 checks for sizing, retries and fallback). - The synthetic fixture generated and packed successfully on CPU: 184 tensors, 15 Q4_0 expert tensors and 5 disk-backed MoE layers. This checks the fixture and pack, not GPU inference. - The CPU image encoder, including its GLM projector, built successfully. The existing llama.cpp MSBuild duplicate-source warning remains. - Changed Python files compiled; JavaScript syntax and Git whitespace checks passed. - CMake rejected GLM with Intel/SYCL and 32-bit Windows. Native Windows x64 GLM configuration reached CUDA discovery and failed with `No CUDA toolset found`. Prepared, but **not executed**: `glm_prefill_parity`, using five synthetic Q4_0 prompt chunks, SSD reads, 12/64 landing slots and read-ahead off/on. It compares 8,192 continuation-logit values with token-at-a-time execution at maximum absolute error below `2e-3`. Minimal GLM Q4_0 fast-kernel/MMQ support is included for this fixture; its CUDA compilation and numerical behavior are not verified here. CUDA compilation/device linking, device parity, actual CUDA pinned-allocation failure/cleanup, and full-model text/image/MTP inference remain pending on compatible NVIDIA/CUDA systems. This machine has no CUDA toolkit; Visual Studio 2026 does not satisfy the recommended CUDA 12.8/Visual Studio 2022 combination. No model weights were downloaded, installed models/configurations/environments were retained, and no production model engine or server was started. The PR remains a draft. Final update commit: `cd2f4a9f540dfa94c06a0823bfacc8d2e7cad355`. ## Strata 0.1.41 conflict resolution — 2026-10-08 Merged official main `fb58e0dbc8399662c0e47c76578c6e878b14f6cf` into this branch without rewriting its published history. Resolved setup CUDA-toolkit-root/GLM flags, baseline-ISA native expert predicates/GLM buffer capacities, and Qwen piece-cache/GLM ignore-merges conflicts. Qwen remains the default. Fresh checks: 38 focused installer, tokenizer, GLM API and request-hardening cases passed; the native-split CPU target rebuilt and passed on MSVC 19.51. These are fresh merge checks, separate from the previous port results. CUDA/device/full-model inference is still pending; no model weights were downloaded or production engine/server started. Conflict-resolution commit: `ff9f6b8832b4d12aba40f8553204d120b7a3cc05`. The newer Maya updates are in draft PR #1476.
本站相关内容
相关页面的快捷入口。