Pull requests / #1499
#1499 core: write reduced helper rows directly to mapped output
open · @W1nge · 0 commentaires · Sur GitHub
Setup & installAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindows
Description
With `--remote-expert-opt`, each active helper layer still reduces into a device buffer and enqueues a device-to-host copy even when the helper already has a mapped output buffer. Write the reduced rows directly to that mapping when `zero_copy_` is available, as the unreduced helper path already does. `finish()` synchronizes the same stream before the host accumulates the rows. The existing device-buffer/copy fallback handles unavailable or disabled mapping. This changes one file (+5/-1), with the same reduction arithmetic and allocations. It uses the existing capability check and adds no hardware names or configuration flags. Validation on Windows WDDM, i7-13850HX / 32 GB RAM, RTX 2080 Ti 22 GiB + P100 16 GiB, driver 537.13, Qwen3.8-Flash-Next IQ3_XXS: - MSVC / CUDA 12.4.131 build passed for `60-real;75-real`, with experimental SM60 enabled. Runtime comparisons used a fixed helper + resident-RAM integration build, changing only this patch for the engine comparisons; the PR itself is based directly on current upstream. - Four long Python code workloads, two rounds: 6,992 output tokens in both builds, **89.41 -> 90.42 token/s (+1.13%)**; complete requests **95.19 -> 94.18 s**. Helper staging/launch time **2.673 -> 2.119 s**, with the same 88,460 layer launches and 4,556 MiB returned. An earlier candidate run was faster; the repeat is the conservative result. These serial runs do not establish a confidence interval or a gain on other hardware. - **57 candidate responses passed** local functional/format checks. The two unchanged-setting coding runs (24 responses) and general run (18 responses) matched baseline text byte-for-byte. Disabling mapping with `STRATA_REMOTE_ZEROCOPY=0` passed three further coding checks, also matching baseline text. - The general spec4 comparison did **not** establish a speedup: complete requests **123.76 -> 123.37 s**, weighted decode **73.86 -> 72.27 token/s (-2.15%)**. I retained the original general deployment and installed the patch only in the local coding configuration. No AMD/HIP hardware was tested. [Full report and retained raw measurements](https://github.com/W1nge/Strata/blob/57e6cbf90e1b76abec51e38c16bfc390dfd72664/docs/performance/PART2-MAPPED-OUTPUT-2026-10-08.md). The local prompt-threshold experiment and its configuration changes are separate from this PR.
Sur le site
Liens install, modèles, releases.