Pull requests / #441
#441 Contrib/non mtp serving
closed · draft · @CC-David-CC · 0 comments · View on GitHub
BenchmarksServer & APIAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Description
Updated to upstream 0.1.34 (`1678de3`). Fresh targeted builds and regressions passed; the full fleet/context performance numbers below remain measurements of the prior 0.1.33 source. [Rebase checks](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/non-mtp-serving/docs/benchmarks/2026-10-02-upstream-sync-serving.md) This is the shared serving split requested in #323 and #336, based on upstream 0.1.34 (`1678de3`). Hardware support and optional performance kernels have separate branches. ## Changes - Allow `--serve` without an MTP drafter. Guard draft binding, prefill, restore, sampling and generation; use one-token verification windows. - Preserve live-prefix reuse. Cross-conversation snapshots still require draft state, so use `--conversation-cache-mib 0` without MTP. - Omit PCIe graph stages only when the expert source has no GPU-visible aliases. A request's `pcie_frac=0` does not remove a path that later requests may need. - Separate prompt-tail and decode profiling samples. - Update the existing identity checker and add a reproducible throughput harness. ## Why it helps On the RX 5500 XT, an 8K-input short code request produced the **same 125 output token IDs** in both modes. Non-MTP completed in **96.44 s**, versus **111.47 s** with MTP: **13.48% less total time**. Prefill was 85.69 versus 103.41 s. Generation alone was slower (11.63 versus 15.52 tok/s). MTP remains better for the measured long RTX PRO and RTX 4090 completions. For example, RTX PRO 64K code produced identical 2213-token outputs at 104.79 / 222.72 tok/s without/with MTP, taking 30.80 / 19.56 s total. The report retains the slower cases and actual output lengths. ## Validation - Six GPU types: RTX 4090, RTX PRO 6000 Blackwell, RTX 3070, one Tesla P4, RX 7900 XTX, and consumer RX 5500 XT 8GB. - Long code/prose at 8K, 32K and 64K input, both modes, up to 4096 output. - Eight identity/replay gates: CUDA, HIP 7, pageable HIP 5.7/7, non-MTP independent-instance replay, and request-level PCIe fraction switching. Each gate uses eight requests per arm, state/output hashes, known answers and live-prefix reuse. The upstream comparison uses MTP; non-MTP replay is identified separately because upstream cannot serve that mode. - RTX 4090 starts at `pcie_frac=0` then alternates 0/0.55. Enabled requests actually perform PCIe work; upstream and candidate tokens/state match. - The initial RTX 3070 64K MTP-on allocation failed. The retained retry with `--vram-reserve-mib 1148` passed both tasks/modes. - RTX 3070 non-MTP full-input test: **131072 input / 2288 output**, 25.28 generation tok/s, 362.25 s prefill, 452.80 s total; 136192 allocated context with Q8 KV streaming. - The separate MTP-on 128K natural-stop request completed at 3893 output tokens; only its exact-length assertion failed. A capacity-only retry with EOS stopping suppressed completed all 4096 output tokens. These results remain distinct. The RX 5500 XT tests use the documented combination of this branch and the hardware contribution. Those source trees and reproduction commands are explicit. [Scope and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/non-mtp-serving/docs/NON_MTP_SERVING_REVIEW.md) | [Full measurements and retained failures](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/non-mtp-serving/docs/benchmarks/2026-10-01-serving.md) | [Machine-readable evidence](https://github.com/CC-David-CC/Strata-a5500/blob/contrib/non-mtp-serving/docs/benchmarks/2026-10-01-serving.json) GSQ-RCO IQ3_S, Q8 KV, greedy text only. Throughput uses MTP4/threshold0.5; identity gates use threshold0. Startup is excluded, the OS file cache is not reset, and outputs may differ. Functional checks are bounded and do not measure general model quality. The maintainer's broader CI was not run locally. Credit: this extends [Niko1221/Strata](https://github.com/Niko1221/Strata); model execution, expert caching and MTP are upstream work.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.