Pull requests / #1001
#1001 generate: wait for the residency table's upload before the next window (#871)
closed · @ischencheng · 0 コメント · GitHub で見る
BenchmarksSetup & installServer & APINVIDIA / CUDAModels & quantsWindowsLinux
本文
On a card that holds every expert, a fresh prompt longer than `--short-read` answers `!!!!` (#871). The batched read lends cache slots to the prompt path (their `host_res` entries become -1), `refill` gives them back, and `res_upload` copies `host_res` to `d_res` with a plain `cudaMemcpy` from a `std::vector`. For pageable memory CUDA only promises the data is staged when that call returns, not that the DMA has landed, and the verify windows run on non-blocking streams that are not ordered after it. If the first zero-doorbell window still sees a -1, `resident_plan_kernel` returns without writing the plan (`skip` is null on that path) and the window computes with an old one. `res_upload` now waits for each copy (`cudaStreamSynchronize(0)` after it). It runs at a lend, a refill, a VRAM resize and an adaptive swap, not per token. Measured on an L40S (49140 MiB, sm_89, the same size as the 48 GB 4090 in the issue), Linux, IQ2_XS, `expert cache 24576 slots` with the zero-doorbell graph, setup's config with `--max-context 32768`. Five fresh prompts over 64 tokens at temperature 0, plus one short one: | engine | `!!!!` replies | |---|---| | main | 0 / 5 | | main, table DMA delayed 50 ms (debug probe) | 5 / 5 | | this PR, same probe | 0 / 5 | | this PR | 0 / 5, same replies as main, same decode rate (220 tok/s) | On Linux the copy lands in time, so main never failed here. The probe (a `cudaLaunchHostFunc` that sleeps 50 ms on stream 0, followed by an async copy of the table) makes the late landing that CUDA allows happen every time. The 41-token prompt was fine in every run, as in the issue. I haven't tested on Windows, where it was reported; my guess is that WDDM batching is what delays the copy there. Side note on #646: on this card the zero-doorbell graph decodes at 216-220 tok/s, against 179 tok/s with `STRATA_VERIFY_ALL_RESIDENT=0` (343-token prompts, 31-36 tokens out). Should fix #871.
関連リンク
インストール・モデル・リリースへの站内リンク。