Pull requests / #314

#314 feat: Qwen3.8-Flash-Next dense model, native Qwen3.5 attention kernel, strata-dense CLI

closed · @perronemirko · 0 评论 · 在 GitHub 查看

Setup & installServer & APINVIDIA / CUDAModels & quantsWindows

描述

## What

Qwen3.8-Flash-Next (dense, Qwen3.5 architecture) as a first-class model path in Strata — no experts, same server protocol as `strata --serve`.

## New

- **`src/core/qwen35.cpp`** (+ header): the Qwen3.5 graph
- **`src/core/dense_model.{cpp,hpp}`**: dense model; weights split across VRAM and pinned host, quantized alpha/beta (Q8_0)
- **`src/kernels/cuda/qwen35_attention.cu`** (+ header): native Qwen3.5 attention kernel — GDN with the SiLU-gated delta norm closure
- **`src/kernels/cuda/dense_kernels.cu`**: dense-path kernels
- **`src/program/dense_main.cpp`**: `strata-dense` executable (new CMake target)

## Changed

- `layer.{cpp,hpp}`: new `g_qwen35_gdn` flag + `layer_set_qwen35_gdn()`; the GDN output-norm path branches to `qwen35_gdn_out_norm`
- `verify.cpp`: the volatile pointer in `raise_flag` moved inside the `_WIN32` guard (was unused on non-Windows)
- `CMakeLists.txt`: new sources in `strata_kernels`/`strata_engine`, `strata-dense` target; dropped the `mmvq_multi_parity` and sampler one-block/old parity tests, simplified two `EXISTS` guards
- Dead code out: `softplus_f` (gdn.cu), double `warp_sum`/`block_sum` (gr.cu), `to_f16_kernel` + f16 include (shared_expert.cu), `CODES_PER_BYTE_S2`, unused helpers in gdn/ple/rope parity
- `.gitignore`: `/data/expert-profile.bin` — regenerate with `tools/make_default_profile.py`

## Tools

- `qwen35_pack.py`, `qwen35_setup.py` — pack + setup for Qwen3.8-Flash-Next
- `dense_setup.py` — dense model setup
- `compare_llama.py` — comparison against llama.cpp
- `dump_mtp.py`, `make_default_profile.py`

## Data

- Added: tokenizer pack under `data/packs/qwen3.8-27b/tokenizer/` (merges, vocab, chat template)
- Removed: committed draft vocabs, expert profiles, experimental speed projection — now generated by the tools above

40 files changed; ~251k insertions (tokenizer merges.txt dominates). 10 commits on `qwen3.8-27b`.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。