Pull requests / #366
#366 [RFC] EXPERIMENTAL Investigate DFlash2 support for Qwen3.8-Flash-Next
closed · draft · @j-luwierski · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Beschreibung
## Summary Related to #347. This draft PR proposes an investigation into **DFlash2 support for Qwen3.8-Flash-Next in Strata**. The goal is to determine whether a compatible DFlash2 drafter can be integrated with Strata’s execution model and provide a measurable improvement over the existing MTP path, while preserving the target model’s decoding semantics. This is initially a planning and feasibility PR. It does **not** claim that a compatible Flash-Next checkpoint is available, that an implementation is complete, or that DFlash2 will be faster on Strata’s supported hardware. I would like to use this PR as a place to discuss the approach, record findings, and keep the investigation reproducible. A documented limitation or negative result would also be a useful outcome. ## Current status and open assumptions The initial investigation separates three questions: 1. **Model availability:** do we have, or can we obtain, a DFlash2 drafter compatible with Qwen3.8-Flash-Next? 2. **Runtime correctness:** can Strata expose the required target features and correctly verify, commit, and roll back draft blocks? 3. **Practical performance:** does the resulting system improve end-to-end generation speed after accounting for drafting, verification, memory pressure, and expert-cache changes? The publicly documented DFlash2 checkpoint for **Qwen3.8-27B** should not be treated as a compatible checkpoint for **Qwen3.8-Flash-Next**. Compatibility needs to be established explicitly. Likewise, Flash-Next’s existing MTP weights are not automatically DFlash2 weights. Reusing parts of the current speculative decoding infrastructure may be possible, but this does not remove the need for a compatible trained drafter. Any statement below describing planned behavior is a proposal, not a completed implementation. ## 1. Model-side requirements This workstream covers the artifacts and architectural contracts required from the target–drafter pair. ### Compatible drafter weights Before implementing the full runtime, establish: - Whether a DFlash2 checkpoint exists for the exact Flash-Next target. - Which target revision and training configuration it was built against. - Whether its architecture, configuration, and reference inference code are available. - Whether the required weights include the DFlash2 selector and dynamic convolution components. - What licenses apply to the checkpoint, reference implementation, and any training data used. A matching tokenizer or similar model name is not sufficient evidence of compatibility. If no compatible checkpoint is available, this becomes a separate training or distillation dependency. The PR should then document the missing requirements rather than present runtime integration alone as working model support. ### Target feature contract Document exactly what the drafter expects from the target: - Source layers and the precise extraction points within those layers. - Whether features are taken before or after normalization or residual operations. - Tensor shapes, ordering, precision, and any concatenation or projection. - Token-to-feature alignment, including the last accepted token and the start of each draft block. - Position IDs, attention masks, and context handling. - Whether and how target embeddings or the output head are reused. - Which features must be retained across successive draft–verify cycles. Layer indices and tensor layouts should come from the checkpoint and reference implementation, rather than being guessed or copied from the 27B integration. ### Token and vocabulary compatibility Verify: - Tokenizer revision and token IDs. - Vocabulary size and ordering. - Special tokens, EOS handling, and any mask or placeholder tokens. - Whether the drafter uses the full vocabulary or a restricted vocabulary. - Whether Strata’s current MTP vocabulary handling can be reused or needs a separate path. Any mismatch should produce an explicit compatibility error. ### If training is required Prepare a separate feasibility note covering: - The intended frozen target checkpoint. - Access to the target features required for training. - Availability of a DFlash2 training implementation. - Data sources, licenses, and train/evaluation separation. - Compute, storage, and memory requirements. - Checkpoint export and conversion requirements. - A small held-out evaluation for draft acceptance and numerical validation. Training a useful drafter should be treated as an independent milestone. It should not be hidden inside the runtime implementation estimate. ## 2. Strata-side changes The runtime work should begin with an audit of existing functionality, especially the MTP verification path, before deciding which components need to change. ### Drafter loading and validation Introduce a clearly defined configuration and loading path for the selected DFlash2 checkpoint. Validation should cover the target identity, architecture, tensor shapes, vocabulary contract, and supported block sizes. Unsupported combinations should fail with actionable diagnostics. The initial implementation should remain experimental and opt-in. ### Target feature extraction Expose the required target features during prefill and decoding. The implementation needs to establish: - Correct feature alignment after full, partial, and zero draft acceptance. - Ownership and lifetime of feature buffers. - Behavior across chunked prefill and subsequent decoding. - The cost of retaining or transferring those features. - Whether extraction changes kernel selection or adds synchronization. Where possible, features should stay on the device that consumes them. Any host transfers should be measured explicitly. ### DFlash2 execution Implement the operations required by the selected reference architecture, including: - Parallel block drafting. - Candidate generation. - The learned path selector. - Dynamic convolution components. - Construction of the final proposed token sequence. The implementation should preserve the reference architecture’s semantics before introducing optimizations. A simpler mock or partial drafter may be useful for testing infrastructure, but it should be clearly identified as such and not reported as DFlash2. ### Verification and state management Audit whether the existing verifier can accept proposals independently of the MTP implementation. Correctness must cover: - Full acceptance, partial acceptance, and rejection at the first candidate. - Committing only accepted tokens and any verifier-produced continuation token. - Restoring attention KV state and recurrent/DeltaNet state after rejection. - Discarding or rebuilding speculative drafter state as required. - Keeping position counters and target feature buffers aligned. - EOS, output limits, and context boundaries. The initial focus should be a linear candidate sequence. More complex verification structures are outside the first milestone unless required by the reference algorithm. ### Sampling semantics Start with greedy decoding to reduce the number of moving parts. Support for nonzero-temperature sampling should be a separate milestone. It requires the correct proposal probabilities and acceptance/correction procedure for the selected DFlash2 implementation. Matching output for one seed is not sufficient evidence of distribution-preserving sampling. Likewise, a deterministic greedy prototype should not be described as supporting general lossless sampling. ### Memory budgeting Account for the complete incremental footprint: - Drafter weights. - Drafter caches and intermediate buffers. - Retained target features. - Candidate and selector buffers. - Verification and rollback scratch space. This is particularly important for Strata because allocating VRAM to a new drafter may reduce the memory available for cached experts. A faster drafting method can still make the complete system slower if it increases CPU work or transfers. ## 3. First small experiment: verifier and rollback correctness The first proposed experiment does not require a trained Flash-Next DFlash2 checkpoint. Its purpose is to establish whether the existing target execution path can safely support an alternative source of draft proposals. ### Setup - One pinned Strata revision. - One supported Flash-Next quantization. - One GPU and one active request. - Short prompts, approximately 2K tokens or less. - Greedy decoding. - No concurrent MTP or other drafting strategy. - Experimental speed projection disabled. - Fixed KV, expert-cache, and numerical execution settings. Use three fixed prompts covering code, prose, and a short reasoning task, generating up to 128 tokens per prompt. ### Procedure 1. Generate a token-by-token reference continuation without speculation. 2. Feed two-token candidate blocks from that continuation into a small verifier test harness. 3. Exercise three cases: - Both candidates match the reference. - The first candidate matches and the second is deliberately incorrect. - The first candidate is deliberately incorrect. 4. After each verification and rollback, continue decoding and compare against the reference. 5. Record the next-token logits, generated token IDs, acceptance decisions, and relevant state positions. The deliberately incorrect token should differ from the reference token and avoid unrelated EOS or special-token behavior. ### Numerical controls Strata’s numerical execution path may depend on token grouping. The experiment must first establish the expected relationship between single-token and multi-token execution under the chosen configuration. Use a controlled kernel configuration where supported. Record tolerances and compare finite logits on identical prefixes before interpreting a later sequence divergence. If the baseline itself is not stable, investigate that first rather than attributing the difference to DFlash2. ### Pass criteria - Correct accepted-prefix lengths in all three cases. - No rejected candidate remains committed in model state. - Subsequent greedy continuation matches the controlled reference. - Next-token logits remain within a predefined numerical tolerance. - No non-finite values are introduced. - State positions and cache lengths remain consistent. Record verification time, rollback time, and additional memory usage, but do not interpret these measurements as DFlash2 performance. **This experiment validates integration infrastructure, not drafter quality or DFlash2 speedup.** ## 4. First experiment with a real drafter Once a compatible checkpoint and reference implementation are available: 1. Validate one drafter forward pass using fixed target features. 2. Compare intermediate tensors, candidate logits, and selected proposals with the reference implementation. 3. Check token/feature alignment across successive cycles. 4. Run a short greedy generation test against target-only decoding. 5. Only after correctness passes, collect acceptance and timing measurements. Start with one supported block size from the checkpoint’s configuration. Any later block-size sweep must distinguish between total block length, number of proposed tokens, and verifier-produced tokens. If the checkpoint is unavailable, this milestone remains blocked; successful mock-proposal tests do not satisfy it. ## 5. Performance evaluation Compare three explicit configurations: | Configuration | Purpose | | --- | --- | | Target-only decoding | Baseline without speculative decoding | | Existing MTP | Current Strata performance reference | | DFlash2 | Proposed integration | Use the same target weights, quantization, prompts, output limits, and sampling settings. Run repeated measurements after warm-up, and report variability rather tha
Mehr auf der Site
Links zu Install, Modellen, Releases.