Pull requests / #1476
#1476 Update GLM port to Project Maya v1.0.4
open · draft · @apr3ndi5 · 0 评论 · 在 GitHub 查看
Setup & installServer & APINVIDIA / CUDAModels & quantsSecurityWindows
描述
## Dependency Depends on #1420, which is still open. This draft is opened against official `main`, so its current diff includes the initial GLM integration from #1420 as well as this update. Review only the new update here: [cd2f4a9...f0fc335](https://github.com/apr3ndi5/Strata/compare/cd2f4a9f540dfa94c06a0823bfacc8d2e7cad355...f0fc335386e51cbe9d8bd3bc951ecc2d08575ff1). After #1420 merges, rebase only the update commit onto the resulting `main` to remove the overlapping baseline. ## Scope and reference Port the changes from Project Maya `5932f601373f53fc021f75dc55159a722c772571` through [`cfd2f45b506be6c02d715e3198b3ff9719de25fc`](https://github.com/mw00/project-maya/commit/cfd2f45b506be6c02d715e3198b3ff9719de25fc). Maya renumbered its releases: the earlier v1.2.0 is v1.0.2, the diagnostic-report update is v1.0.3, and the tensor-core/FP16-cache update is v1.0.4. Strata's own version is unchanged. ## Changes - Port prompt MLA attention to tensor cores with FP16 operands, F32 accumulation and next-cell prefetching. Preserve `STRATA_GLM_PREFILL_ATTN=f32` and `STRATA_GLM_PREFILL_ATTN_CHECK` for comparison. - Store the fast engine's DSA latent cache in FP16 across prefill, token decoding and MTP. The diagnostic SLOW path retains F32. This halves latent-cache storage, not the whole engine's memory. - Add `--family glm --report` to the existing installer. It writes local hardware, available RAM/commit, recorded build information, settings and recognized numeric engine metrics to `strata-glm-report.txt`. It does not install, download, start or upload anything. Keys, conversations, arbitrary log tails and personal paths are excluded. - Add prepared attention parity tests and extend the Q4_0 prefill test to F32/automatic attention selection, 12/64 landing slots and read-ahead off/on. Preserve the existing `2e-3` numerical tolerance. Strata compatibility adaptations align packed cache starts to 16 bytes, round storage up, guard WMMA on pre-sm_70 builds, and reduce F32 fallback tiles to fit below 48 KiB of shared memory on Turing. Qwen remains the default; model hashes, APIs, dashboard, licenses, installed configurations and external-host API-key protection are retained. ## Fresh validation for this update Run locally on Windows on 2026-10-07: - **596 Python cases completed: 589 passed, 7 skipped.** Mock server/API/frontend/security/lifecycle/monitor/VRAM/detokenization: 331 passed, 5 skipped. Installer/pack/tokenizer/oracle/configuration/responses/Qwen regression: 258 passed, 2 skipped. The 22 GLM installer/report cases include missing tools/corrupt state, bounded log reads, redaction, unsupported platforms, download consent and no-start behavior. Reruns are not counted twice. - **4 C++ CPU tests passed:** conversation cache, conversation memory, native split, and GLM prefill-memory helper. Its 19 checks include FP16 storage/alignment and unchanged diagnostic F32 layout. - Synthetic Q4_0 fixture generation/packing passed on CPU: 184 tensors, 15 Q4_0 expert tensors, 5 disk-backed layers. This is not inference validation. - CPU image encoder, including the GLM projector, built with MSVC 19.51. The existing llama.cpp duplicate-source MSBuild warning remains. - Python/JavaScript syntax and Git whitespace checks passed. - Native Windows GLM CUDA configuration reached toolkit discovery and stopped with `No CUDA toolset found`. Prepared but **not compiled or run on a device**: `glm_batch_attention_f32` and `glm_batch_attention_wmma` compare 81,920 context values per mode with a CPU softmax reference and check FP16 guards, negative/empty cells and partial tiles. `glm_prefill_parity` compares 16,384 continuation logits across eight configurations with streaming execution. CUDA compilation/device linking, tensor-core parity, Turing fallback, full-model quality, images and one/two-GPU MTP remain pending on compatible NVIDIA/CUDA hardware. This checkout has no CUDA toolkit and has Visual Studio 2026; CUDA 12.8 with Visual Studio 2022 is still the recommended Windows combination. No performance numbers from Maya are presented as measurements of Strata. No weights were downloaded and no production model engine or server was started. ## Commits Head: `apr3ndi5:codex/glm5-maya-v1.0.4`. Update commit: `f0fc335386e51cbe9d8bd3bc951ecc2d08575ff1`. Parent: `cd2f4a9f540dfa94c06a0823bfacc8d2e7cad355` from #1420. Keep this PR in draft until compatible CUDA compilation and device validation are available.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。