Pull requests / #943

#943 Add experimental deepMoE Vulkan backend for DeepSeek text chat

closed · @cklxx · 0 comentarios · En GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationLinux

Descripción

Strata's existing engine boundary can also drive a native DeepSeek process. This adds an explicit `--engine deepmoe` backend for [deepMoE](https://github.com/acupof-ai/cachedMoE), a C++20 Vulkan/Slang engine that streams the native FP4/FP8 DeepSeek-V4.1-Flash checkpoint through an NVMe expert cache on a 128 GB AMD Strix Halo PC.

The same Strata web app and OpenAI/Anthropic/Responses service paths can serve text chat. The adapter uses the checkpoint's own prompt renderer, a CPU tokenizer, native EOS IDs, incremental UTF-8 output, progress events, persistent subprocess reuse, and cancellation with drain before the next request. The web app disables unsupported top-k sampling for this backend. Existing Qwen engines and installer behaviour remain the defaults.

Scope is deliberately experimental and explicit: source-built Linux/RADV, text chat, one engine. DeepSeek DSML tools, images, MCP, parallel slots, lazy loading, and the one-click installer are not integrated. Unsupported controls are rejected. This does not port DeepSeek kernels into the CUDA/HIP engine, claim compatibility with Qwen's hardware requirements, or enable approximate mask/speculative modes. `docs/DEEPMOE.md` has setup, a config example, supported controls, and quality/performance links.

Validation:
- `python -m unittest serve.test_server serve.test_detok serve.test_deepmoe`: 158 tests, passed; 5 environment-dependent skips.
- `DEEPMOE_MODEL_DIR=... python -m unittest serve.test_deepmoe_native`: 2 CPU checkpoint checks passed (tokenization, split UTF-8, ordinary quoted thinking tags, native prompt/effort).
- JavaScript syntax check and diff whitespace check passed.
- The fake-process tests cover consecutive turns, early iterator close, cancellation, process restart, EOS, unsupported sampling, OpenAI streaming/non-streaming, and Anthropic output.

Real-process smoke passed on Linux/RADV with an AMD Radeon 8060S (128 GB Strix Halo), two checkpoint read sources, 5,000 expert slots, exact routing, and a 4K context capacity. OpenAI chat returned “2 + 3 equals 5.”, Anthropic chat returned “3 + 4 equals 7.”, and OpenAI streaming returned a separate thought process followed by “4 + 5 equals 9.” The native HTML identifies this backend so the web app uses supported sampling defaults. A fresh engine also restored a saved 75-token KV snapshot and answered successfully. These are interoperability checks, not a throughput or quality benchmark. [Hardware/KV receipt](https://github.com/acupof-ai/cachedMoE/blob/main/docs/kv_async_receipt.json).

The initial hardware attempt caught Qwen's config validator rejecting native top-k=0. Native sampling is now validated before starting the subprocess, with a main-entry regression test and successful hardware rerun. The API integration test also covers a basic Responses request. There is no new throughput claim.

En el sitio

Enlaces a install, modelos, releases.