Pull requests / #1043

#1043 generate: residency-table uploads wait for their own copy before the non-blocking streams read the table

closed · @sergqwer · 0 commentaires · Sur GitHub

Server & APINVIDIA / CUDA

Description

Replaces #550. GitHub closed it on 2026-10-05, when my fork was made private by mistake: that took the fork out of the network for good, so it can no longer open pull requests. This is the same branch and commit (`000f288`), opened from a new fork; the discussion and the measurements are in #550.

---

Every upload of the expert residency table, host_res to `d_res` and each stage's, is a plain `cudaMemcpy` (`src/program/generate.cpp`, 9 call sites). The kernels that read the table run on streams created `cudaStreamNonBlocking`.

- A pageable host-to-device `cudaMemcpy` may return once the data is staged, before its DMA has landed.
- A non-blocking stream does not wait for the legacy stream the copy runs on.

So nothing orders a kernel enqueued right after an upload (the next verify window after an adaptive swap, a trim, a refill) behind the copy. That kernel could read a table half old, half new: an expert computed by both the GPU and the CPU, or by neither.

The table is ~100 KB and usually lands in microseconds, so this rarely fires. It is the same class as #532 and #536, where the tests' uploads raced their non-blocking streams. Ours made one of those tests fail in 6 of 20 runs.

**Fix:** the uploads go through `res_put()`, which waits for the copy on the current device's legacy stream (`cudaStreamSynchronize(cudaStreamLegacy)`): that copy only, not the device.

**Measured** on an RTX 5090, IQ2_XS (ISTA), `--expert-cache 14900`, v0.1.37 against this PR:
- First-token logits after a 2K and a 32K prompt: byte-identical.
- Decode (a 32K prompt, then up to 600 tokens, `--stop-eos`), 6 interleaved pairs: 15.82 ms a round against 15.93 (+0.7%, within noise). One more pair is left out: 19.80 ms on a different token path, a desktop stall.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

Sur le site

Liens install, modèles, releases.