Pull requests / #1524
#1524 Performance: wider speculative batches and shared experts — 1.31×–2.07× c4 throughput
open · @tntcannon5000 · 0 评论 · 在 GitHub 查看
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
描述
Supersedes #1209 with a performance-focused port onto **pristine upstream v0.1.40.4 (`6674a00`)**. Wider speculative batches let concurrent requests share more verification and expert work while retaining upstream admission, prefix reuse, cancellation, checkpoints and return-to-solo handling. Fresh c=4 measurements on an **RTX 5090, Ryzen 9 9950X3D, 48 GB DDR5-6000** show **1.31x essay and 2.07x coding aggregate decode throughput**. These replace the older PR's figures; both arms were rebuilt and measured again. ## Performance against current upstream Three paired runs per workload, alternating AB/BA/AB, seeds 123/456/789. Values are medians; ratios are ratios of medians. Baseline has upstream concurrent MTP enabled (`--batch-mtp`) and needs no correctness patch. | Workload | Measurement | Upstream TPS | Candidate TPS | Gain | | --- | --- | ---: | ---: | ---: | | Essay writing | Four-stream decode | 154.8 | 203.1 | 1.31x | | Essay writing | Whole bounded run | 149.6 | 190.5 | 1.27x | | Tool-driven coding | Four-stream decode | 162.5 | 336.0 | 2.07x | | Tool-driven coding | Whole bounded run | 136.3 | 233.5 | 1.71x | Swift 1.5 IQ2_XS; 98,304 allocated context/request; int8 KV; 2,560 MiB requested reserve; identical explicit cache request, rounded by profile sizing to **11,568 resident experts in every run**. Temperature .85, normal optimized CPU arithmetic, QFUSE off. Fresh processes with startup and 256-token warm-up excluded. No cache cuts or reserve shrinks occurred. GPU undervolt removal was previously reported by the owner; the voltage curve was not independently verified. Median native token-gap p95: **47 -> 47 ms for essays; 62 -> 47 ms for coding**. Turn TTFT p95: .313 -> .344 s and .891 -> .813 s respectively. These are native receipt measurements, not browser latency guarantees. Four-stream decode counts emitted tokens while all four slots are active with no admission. Whole-run throughput includes prompt work, tools and periods with fewer streams. The coding fixture is a bounded read/write/AST-check loop, not DSH or SWE-bench. **12/12 modules passed syntax/structure checks in each arm**, but nine agents per arm reached the six-turn limit. Generated tests were not executed. Outputs differ, so these numbers do not establish faster equal-quality task completion. [Full method and results](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/PR_WORKLOAD_BENCHMARKS.md) | [All twelve run summaries](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/evidence/v0404-performance/summary.json) ## What changes - Wider target verification across 2-4 concurrent streams, with 8/12/16-row capacities and accepted-prefix state commits. - Per-stream MTP and suffix proposals, shared draft weights with private stream state, optional overlapping drafting, stable layouts and bounded graph caching. - Merged expert execution, optional parallel dense groups, wider CPU expert processing, and optional residency/prefill/fairness/PCIe controls. The speedup is for the combined opt-in configuration, not an isolated attribution to each feature. Wider speculation remains **off by default** and limited to eligible single-NVIDIA, text/MTP, resident-int8-KV configurations without helper expert caches. Unsupported configurations keep upstream batching; AMD/multi-GPU concurrency is not disabled. ## How this differs from #1209 - **No Python server or prefill telemetry changes.** Personal launchers and configs are also excluded. - Keep upstream's batched QFUSE correction and shared-MTP-head type fix instead of duplicating them. The old PR's remaining QFUSE arithmetic/solo-commit changes are not included or claimed fixed here. - Preserve upstream's new interleaved MMVQ, PDL, dense-branch, Q4-query and PLE paths. Parallel groups receive separate interleaved scratch; wide output projection is tiled to supported kernel widths. - Keep solo graph arrays bounded to eight rows independently of wider batch capacity. - Exclude the separate c8/24/32-row experiments. [Integration decisions and audit](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/V0404_REBASE.md) | [Feature rationale](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/CONCURRENT_THROUGHPUT.md) ## Fresh validation and limits - Native tests passed: **9,840 layouts**, **100 bitwise partial commits across 16 layouts**, and the sixteen-row IQ2_XS/Q2_0 expert fixture with **zero failures**. - Controlled four-stream/solo comparisons passed on both builds. Cross-build dumps matched **512 solo + 512 batch tokens** exactly. - **All eight request-lifecycle comparisons passed on both builds**, including prompt/decode interleaving, yield/resume, slot continuation, checkpoint reuse and return to solo. - Exact-token controls fix expert placement and CPU arithmetic; they are separate from the normal fast-mode benchmark. Normal adaptive token identity, universal quality equivalence and full-window responsiveness are not claimed. - Default interleaved MMVQ is exercised. QFUSE-on, additional upstream opt-in PDL/dense-branch combinations, AMD, Intel and multi-GPU hardware are not separately qualified. No new c=1 speedup is claimed. [Qualification summary, logs, token dumps and artifact hashes](https://github.com/tntcannon5000/Strata/tree/publish/concurrent-performance-v0404/docs/evidence/v0404-performance)
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。