Pull requests / #1524

#1524 Performance: wider speculative batches and shared experts — 1.31×–2.07× c4 throughput

open · @tntcannon5000 · 0 comments · View on GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Description

Supersedes #1209 with a performance-focused port onto **pristine upstream v0.1.40.4 (`6674a00`)**. Wider speculative batches let concurrent requests share more verification and expert work while retaining upstream admission, prefix reuse, cancellation, checkpoints and return-to-solo handling.

Fresh c=4 measurements on an **RTX 5090, Ryzen 9 9950X3D, 48 GB DDR5-6000** show **1.31x essay and 2.07x coding aggregate decode throughput**. These replace the older PR's figures; both arms were rebuilt and measured again.

## Performance against current upstream

Three paired runs per workload, alternating AB/BA/AB, seeds 123/456/789. Values are medians; ratios are ratios of medians. Baseline has upstream concurrent MTP enabled (`--batch-mtp`) and needs no correctness patch.

| Workload | Measurement | Upstream TPS | Candidate TPS | Gain |
| --- | --- | ---: | ---: | ---: |
| Essay writing | Four-stream decode | 154.8 | 203.1 | 1.31x |
| Essay writing | Whole bounded run | 149.6 | 190.5 | 1.27x |
| Tool-driven coding | Four-stream decode | 162.5 | 336.0 | 2.07x |
| Tool-driven coding | Whole bounded run | 136.3 | 233.5 | 1.71x |

Swift 1.5 IQ2_XS; 98,304 allocated context/request; int8 KV; 2,560 MiB requested reserve; identical explicit cache request, rounded by profile sizing to **11,568 resident experts in every run**. Temperature .85, normal optimized CPU arithmetic, QFUSE off. Fresh processes with startup and 256-token warm-up excluded. No cache cuts or reserve shrinks occurred. GPU undervolt removal was previously reported by the owner; the voltage curve was not independently verified.

Median native token-gap p95: **47 -> 47 ms for essays; 62 -> 47 ms for coding**. Turn TTFT p95: .313 -> .344 s and .891 -> .813 s respectively. These are native receipt measurements, not browser latency guarantees.

Four-stream decode counts emitted tokens while all four slots are active with no admission. Whole-run throughput includes prompt work, tools and periods with fewer streams. The coding fixture is a bounded read/write/AST-check loop, not DSH or SWE-bench. **12/12 modules passed syntax/structure checks in each arm**, but nine agents per arm reached the six-turn limit. Generated tests were not executed. Outputs differ, so these numbers do not establish faster equal-quality task completion.

[Full method and results](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/PR_WORKLOAD_BENCHMARKS.md) | [All twelve run summaries](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/evidence/v0404-performance/summary.json)

## What changes

- Wider target verification across 2-4 concurrent streams, with 8/12/16-row capacities and accepted-prefix state commits.
- Per-stream MTP and suffix proposals, shared draft weights with private stream state, optional overlapping drafting, stable layouts and bounded graph caching.
- Merged expert execution, optional parallel dense groups, wider CPU expert processing, and optional residency/prefill/fairness/PCIe controls.

The speedup is for the combined opt-in configuration, not an isolated attribution to each feature. Wider speculation remains **off by default** and limited to eligible single-NVIDIA, text/MTP, resident-int8-KV configurations without helper expert caches. Unsupported configurations keep upstream batching; AMD/multi-GPU concurrency is not disabled.

## How this differs from #1209

- **No Python server or prefill telemetry changes.** Personal launchers and configs are also excluded.
- Keep upstream's batched QFUSE correction and shared-MTP-head type fix instead of duplicating them. The old PR's remaining QFUSE arithmetic/solo-commit changes are not included or claimed fixed here.
- Preserve upstream's new interleaved MMVQ, PDL, dense-branch, Q4-query and PLE paths. Parallel groups receive separate interleaved scratch; wide output projection is tiled to supported kernel widths.
- Keep solo graph arrays bounded to eight rows independently of wider batch capacity.
- Exclude the separate c8/24/32-row experiments.

[Integration decisions and audit](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/V0404_REBASE.md) | [Feature rationale](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance-v0404/docs/CONCURRENT_THROUGHPUT.md)

## Fresh validation and limits

- Native tests passed: **9,840 layouts**, **100 bitwise partial commits across 16 layouts**, and the sixteen-row IQ2_XS/Q2_0 expert fixture with **zero failures**.
- Controlled four-stream/solo comparisons passed on both builds. Cross-build dumps matched **512 solo + 512 batch tokens** exactly.
- **All eight request-lifecycle comparisons passed on both builds**, including prompt/decode interleaving, yield/resume, slot continuation, checkpoint reuse and return to solo.
- Exact-token controls fix expert placement and CPU arithmetic; they are separate from the normal fast-mode benchmark. Normal adaptive token identity, universal quality equivalence and full-window responsiveness are not claimed.
- Default interleaved MMVQ is exercised. QFUSE-on, additional upstream opt-in PDL/dense-branch combinations, AMD, Intel and multi-GPU hardware are not separately qualified. No new c=1 speedup is claimed.

[Qualification summary, logs, token dumps and artifact hashes](https://github.com/tntcannon5000/Strata/tree/publish/concurrent-performance-v0404/docs/evidence/v0404-performance)

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.