Pull requests / #807

#807 cuda: opt-in batched expert uploads to reduce host submission overhead

closed · draft · @midhatn · 0 commentaires · Sur GitHub

Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux

Description

The verifier's DMA path currently submits one copy per expert. This adds an opt-in `cudaMemcpyBatchAsync` path for independent pinned expert uploads on CUDA 13+, using the existing copy stream and readiness callback.

The useful result is lower host submission overhead. It is **not a demonstrated model-speed improvement**, so this is a draft and the default stays off.

On Windows with a Ryzen 9 7940HS, RTX 4070 Laptop 8 GB and 64 GB DDR5-5600, 16 registered-host uploads of 2 MiB each measured:

| Path | Median submission wall time | Median total transfer time |
| --- | ---: | ---: |
| Individual copies | 0.0936 ms | 2.5589 ms |
| Batched copies | 0.0162 ms | 2.5248 ms |

That is 82.7% lower submission latency but only 1.3% lower total transfer latency. A separate repetition using `QueryThreadCycleTime` found about 84.9% fewer submitting-thread cycles. These are separate measurements; neither means an equivalent reduction in whole-model CPU usage or task time.

`STRATA_DMA_BATCH=1` enables batching with `--pcie-mode dma`. Mode `2` additionally requests CUDA 13.4's overlap-with-compute hint. Older CUDA and HIP builds retain individual copies, as do single uploads and groups larger than 128. The normal kernel-upload path is unchanged. Registered Windows host pointers are resolved to device-visible aliases, and submission errors fail the engine rather than releasing consumers with stale weights.

Validation:

- CUDA 13.4 and 13.0 batch paths, plus the CUDA 12.6 fallback: 63 byte-parity/event-order cases for each of `cudaMallocHost` and Windows `VirtualAlloc` + `cudaHostRegister`, 378 cases in total. The CMake test targets also pass.
- The isolated patch on v0.1.39 passed 53 recorded model requests plus an HTTP smoke check: high reasoning, Swift 1.5 IQ3_XXS, 64K allocated context, MTP 5, tools, prompt reuse, a 57,750-token prompt with auxiliary-call restoration, and CPU vision.
- A-B-B-A coding runs passed independent checks but had different reasoning lengths and draft acceptance. They did not establish better time to a verified answer. The table and settings are in `docs/BATCHED_DMA.md`.
- Linux and HIP runtime testing remains outstanding. The older-CUDA fallback was tested on Windows, not HIP.

The design reference is [NVIDIA NCCL's independent-copy batching](https://github.com/NVIDIA/nccl/blob/12df1a11afad322be5a204a2db890161cbf8131d/src/ce_coll.cc), using the [CUDA batch-copy contract](https://docs.nvidia.com/cuda/cuda-runtime-api/cuda_runtime_api/group__CUDART__MEMORY.html). The helper and tests are new; no NCCL source was transplanted. No other experimental Strata PR is included.

Sur le site

Liens install, modèles, releases.