Pull requests / #885

#885 Perf/duplex transfers only

closed · @CC-David-CC · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

## Result and scope

On an RTX PRO 6000 Blackwell 96 GB with Q8_0, FP16 KV and MTP T4, the locked-table coding comparison improved decode from 129.21 to 136.79 tok/s (+5.87% across two reversed-order pairs). The individual gains were +5.35% and +6.39%. All output token IDs and recorded work matched.

There was no consistent full-request improvement: this workload was dominated by prefill. These are small-sample results for the documented resident-RAM configuration, not a preset-wide speedup claim.

## Change

`STRATA_EXCHANGE_DUPLEX=1` overlaps resident expert evictions and refills on separate CUDA streams. Each refill waits for its own slot's eviction event; existing completion and ownership barriers remain in place. This changes scheduling, not transfer bytes, model quantization or expert selection.

- Off by default; existing presets unchanged.
- Explicit resident RAM is required: `--resident-budget-gib N` or `--resident-experts`.
- Single-GPU serving with pinned buffers and a fixed cache is supported. Unsupported configurations retain sequential copies.
- No dependency on exchange-buffer rotation. Q8 integration evidence separately includes the Q8 reader and rotation.
- Base: current upstream `6f32ec0`. Three focused commits; large raw evidence stays on the validation branch.

## Validation

- Component fixtures passed on RTX PRO 6000 and RTX 4090.
- Byte checks cover different blob and slot sizes, invalid inputs, repeated AB/BA orders, close/reopen and early teardown with transfers pending.
- CUDA memcheck: zero errors. ASan/UBSan fixture passed with the documented CUDA-compatible settings; leak detection was not run.
- 37 exact-token/recorded-work comparisons across 76 Q2/Q8 requests, including target-only and MTP, default-off and fallback paths, cancellation/recovery, and an actual 128K-input case.
- Q2 MTP decode improved about 1.7-2.8% in the repeated coding/editing cases. Target-only 32K did not improve.

The earlier unlocked-table Q8 timings were confounded and are not used as the speed claim. The locked-table repeats still recorded some system-wide swap-ins/major faults; limitations and all request times are retained in the report.

[Full measured results, ordering contract and reproduction](https://github.com/CC-David-CC/Strata-a5500/blob/perf/duplex-transfers-only/docs/DUPLEX_EXCHANGES.md)

[Raw evidence and correctness checks](https://github.com/CC-David-CC/Strata-a5500/tree/test/duplex-q8-integration/bench/results/2026-10-05-duplex)

Mehr auf der Site

Links zu Install, Modellen, Releases.