Pull requests / #1209

#1209 Speed up concurrent generation, fix QFUSE, and improve prefill monitoring

open · @tntcannon5000 · 0 Kommentare · Auf GitHub

BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

This adds opt-in wider speculative decoding across concurrent requests, fixes QFUSE rounding/state bugs, and makes prefill monitoring reflect actual prompt work. It preserves upstream request handling and keeps the existing paths for configurations that cannot use the wider batches.

On an **undervolted RTX 5090, Ryzen 9 9950X3D, 48 GB DDR5-6000**, median aggregate decode throughput at c=4 increased **1.37x for essay writing** and **1.95x for tool-driven coding** in three paired runs per workload.

## What changed and why

1. **Concurrent throughput:** verify more proposed tokens together to reuse weights and GPU work across requests. Includes per-slot MTP and suffix proposals, private draft state with shared weights, accepted-prefix commits, 8/12/16-row capacities, optional draft overlap and merged expert execution, wider CPU expert processing, and optional scheduling controls. Upstream admission, prefix reuse, cancellation, preemption, checkpoints and return-to-solo behavior remain in place.
2. **QFUSE correctness:** match the consuming quantizer's finite-half guards and division/rounding convention, retain required batched GDN quantization, and align one-token commit eligibility with capture behavior. Global math flags and defaults are unchanged.
3. **Prefill monitoring:** emit explicit engine prompt-phase metadata so queued requests and stale prefill rates are not reported as ongoing prompt work. Older engines retain the fallback behavior. This part changes observation, not scheduling or sampling.

Wide speculation is **off by default** and applies to eligible 2–4-slot, single-NVIDIA, text/MTP, resident-int8-KV configurations without helper expert caches. Other configurations retain upstream batching. Personal launchers, model configs and daily-use server modifications are excluded; the Python server diff is the scoped monitoring change.

## Performance

Swift 1.5 IQ2_XS, four concurrent requests, 98,304 context capacity per request, int8 KV, 2,560 MiB reserve; both arms loaded 11,022 experts. Temperature .85; normal optimized CPU kernels and adaptation enabled; thinking disabled for this bounded fixture. Each run uses a fresh process and excludes warm-up.

| Workload | Measurement | Baseline TPS | Candidate TPS | Ratio |
| --- | --- | ---: | ---: | ---: |
| Essay writing | Four-stream decode | 146.2 | 200.1 | 1.37x |
| Essay writing | Whole bounded run | 141.1 | 189.2 | 1.34x |
| Tool-driven coding | Four-stream decode | 167.1 | 325.7 | 1.95x |
| Tool-driven coding | Whole bounded run | 137.3 | 221.5 | 1.61x |

**Baseline qualification:** the reference is upstream v0.1.40.1 (`82f46a8`) plus the two-line shared-MTP-head metadata fix included here (`dhead_type_` and `dvocab_host_`). Pristine upstream failed concurrent MTP admission with `unsupported native MMVQ GGML type`. Both measured arms therefore have working concurrent MTP; this is not a pristine-binary comparison.

Four-stream decode measures emitted tokens during intervals with four active slots and no admission underway. Whole-run throughput includes prompt reads, tools and periods with fewer active streams. Outputs differ, so elapsed time is not an equal-work completion comparison.

The coding fixture uses four independent read/write/check tool loops, not DSH or SWE-bench. Checks compile syntax and inspect structure; generated code/tests are not executed. **12/12 baseline modules and 11/12 candidate modules passed structure checks.** Ten baseline agents and nine candidate agents hit the six-turn cap; only two and three respectively finished naturally with a passing checked module. These results establish throughput on these fixtures, not faster correct task completion or equivalent quality.

[Full benchmark method, configuration and exclusions](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance/docs/PR_WORKLOAD_BENCHMARKS.md) · [All 12 run summaries](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance/docs/evidence/pr-workloads-summary.json)

## Validation and limitations

- Fresh pre-publication run: **64 server/setup tests passed**.
- Fresh runs of **six native tests passed**: batch layouts (9,840 layouts), partial state commits (100 commits across 16 layouts), expert multi-row processing, GR parity, QFUSE quantization boundaries and QFUSE GDN.
- Native checks reused the previously built integrated binaries. The combined PR's runtime source is identical to that tested candidate; hashes and logs are included. No new full build is claimed for the packaging step.
- Historical controlled QFUSE off/on generation matched **7,638/7,638 tokens**. The isolated quantization correction restored **150/150 full-vocabulary positions byte-for-byte**. These controls fixed expert placement and CPU arithmetic.
- Normal adaptive outputs remain non-identical. Historical comparisons against the corrected older v039 candidate passed the recorded absolute KL/TV screens, while the stricter .0001 mean-KL trigger still flags differences, including self-repeats. Those comparisons are not a pristine-upstream quality benchmark.
- AMD/multi-GPU fallback paths are preserved but lack hardware qualification. No new c=1 speedup, universal quality equivalence, or full-window responsiveness guarantee is claimed.

[Validation evidence and reproduction](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance/docs/PR_QUALIFICATION.md) · [Performance change audit](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance/docs/CONCURRENT_THROUGHPUT.md) · [QFUSE rationale](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance/docs/QFUSE_CORRECTNESS.md) · [Monitoring protocol](https://github.com/tntcannon5000/Strata/blob/publish/concurrent-performance/docs/PREFILL_MONITORING.md)

The commits separate correctness, performance, benchmark documentation, monitoring, and final evidence for review and rollback. This combines the three originally prepared changes into one submission.

Mehr auf der Site

Links zu Install, Modellen, Releases.