Pull requests / #807
#807 cuda: opt-in batched expert uploads to reduce host submission overhead
closed · draft · @midhatn · 0 comentarios · En GitHub
Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descripción
The verifier's DMA path currently submits one copy per expert. This adds an opt-in `cudaMemcpyBatchAsync` path for independent pinned expert uploads on CUDA 13+, using the existing copy stream and readiness callback. The useful result is lower host submission overhead. It is **not a demonstrated model-speed improvement**, so this is a draft and the default stays off. On Windows with a Ryzen 9 7940HS, RTX 4070 Laptop 8 GB and 64 GB DDR5-5600, 16 registered-host uploads of 2 MiB each measured: | Path | Median submission wall time | Median total transfer time | | --- | ---: | ---: | | Individual copies | 0.0936 ms | 2.5589 ms | | Batched copies | 0.0162 ms | 2.5248 ms | That is 82.7% lower submission latency but only 1.3% lower total transfer latency. A separate repetition using `QueryThreadCycleTime` found about 84.9% fewer submitting-thread cycles. These are separate measurements; neither means an equivalent reduction in whole-model CPU usage or task time. `STRATA_DMA_BATCH=1` enables batching with `--pcie-mode dma`. Mode `2` additionally requests CUDA 13.4's overlap-with-compute hint. Older CUDA and HIP builds retain individual copies, as do single uploads and groups larger than 128. The normal kernel-upload path is unchanged. Registered Windows host pointers are resolved to device-visible aliases, and submission errors fail the engine rather than releasing consumers with stale weights. Validation: - CUDA 13.4 and 13.0 batch paths, plus the CUDA 12.6 fallback: 63 byte-parity/event-order cases for each of `cudaMallocHost` and Windows `VirtualAlloc` + `cudaHostRegister`, 378 cases in total. The CMake test targets also pass. - The isolated patch on v0.1.39 passed 53 recorded model requests plus an HTTP smoke check: high reasoning, Swift 1.5 IQ3_XXS, 64K allocated context, MTP 5, tools, prompt reuse, a 57,750-token prompt with auxiliary-call restoration, and CPU vision. - A-B-B-A coding runs passed independent checks but had different reasoning lengths and draft acceptance. They did not establish better time to a verified answer. The table and settings are in `docs/BATCHED_DMA.md`. - Linux and HIP runtime testing remains outstanding. The older-CUDA fallback was tested on Windows, not HIP. The design reference is [NVIDIA NCCL's independent-copy batching](https://github.com/NVIDIA/nccl/blob/12df1a11afad322be5a204a2db890161cbf8131d/src/ce_coll.cc), using the [CUDA batch-copy contract](https://docs.nvidia.com/cuda/cuda-runtime-api/cuda_runtime_api/group__CUDART__MEMORY.html). The helper and tests are new; no NCCL source was transplanted. No other experimental Strata PR is included.
En el sitio
Enlaces a install, modelos, releases.