Pull requests / #1315

#1315 glm5-next: GLM-5.3-Flash support {intial attempt & prefill really sucks avoid testing its not ready }

open · @gopinath87607 · 0 评论 · 在 GitHub 查看

BenchmarksNVIDIA / CUDA

描述

GLM-5.3-Flash is a 312B-total / 17B-active hybrid MoE: 45 trunk layers, 11 absorbed-NoPE-MLA full-attention layers (every fourth) and 34 KDA linear-attention layers, all wrapped in hyper-connections.

Three things in a glm5-next block have no analogue in this engine, and each one produces plausible output when wrong rather than crashing:

  * HYPER-CONNECTIONS own every residual.  The sublayer's output is mixed back into four streams through a Sinkhorn-normalised 4x4 matrix computed per token from a 24-wide projection with its own base and scale.  Ported from the reference's ggml_hc_pre/ggml_hc_post, not re-derived.
  * KDA is a delta-rule recurrence with a per-head state, a depthwise causal convolution over three SEPARATE q/k/v streams that nevertheless share one state (the reference concatenates their filters before convolving), an L2 normalisation whose eps is a FLOOR on the norm rather than a term inside the sum, and a decay gate bounded to (-5, 0).
  * ABSORBED NoPE MLA.  K and V are the same tensor: the cache holds the 512-wide kv_lora latent, one row per token, and one token's value IS that latent.  There is one KV head, no RoPE (rope.dimension_count is 0), and the softmax scale is 1/sqrt(256) - the per-head width BEFORE absorption - not 1/sqrt(512).

The router is not qwen4exp's either: sigmoid, with no softmax anywhere in the model, a bias that steers SELECTION only, and weights that are the UN-BIASED probabilities normalised by their SUM. ffn_gate_inp.weight is F32 here where qwen4exp's is BF16 - reading it as BF16 gives 288 plausible logits and a different top-8.

New: five kernel files (glm_mhc, glm_kda, glm_delta, glm_attn, glm_elt), the layer order and its buffer carve (glm_layer.{hpp,cpp}), expert-pool binding (glm_experts.{hpp,cpp}), architecture detection (model_arch.cpp, arch.hpp), the projection helper layer.cpp and glm_layer.cpp now share (gemv_util.{hpp,cpp}), and glm_parity - a ctest in which every case also computes a RIVAL reading and requires the two to be told apart, because each of these choices yields well-formed output when it is wrong.

Measured on 1x RTX 3060 12 GB, UD-IQ4_XS 5 shards, 319-token prompt, greedy: prefill 6.42 tok/s at --prefill 1, 9.81 at 128 and 10.41 at 4096; decode 6.36 tok/s.  The greedy ids are identical at all three chunk sizes and equal to the reference captured before this work landed.

Not included: vision, MTP, speculative decoding, --batch and --gpu-stages stay out for this arch.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。