反馈 / #669
#669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores
open · @gputier · 4 评论 · 去 GitHub 看
BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows
说明
Measurements of the prompt path at a 1M context on an RTX 5090, with what I tried and what did not help. No request attached: two observations you may want to act on, and three dead ends so that nobody repeats them. ## Setup Windows 11, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB of RAM. Engine built from v0.1.38 (99f3dbd), Release, sm_120, CUDA 13.4, MSVC 14.44. Qwen3.8-Flash-Next GSQ-RCO IQ3_S, served by `serve/server.py` over `/v1/messages`. Engine arguments: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 1048576 --kv int8 --kv-resident 32768 --rope-scaling yarn --rope-scale 4 --conversation-cache-mib 16384 --conversation-cache-slots 4`. The expert cache has 11,702 slots and 205 MiB of VRAM stay free, in every run below. Two prompts, always in this order after a fresh start: 58,342 tokens read from 0, then 597,889 tokens whose first tokens are the first prompt (so it resumes from a checkpoint). Synthetic log lines, one question about a line in the middle; the answer was right in every run. ## 1. `--prefill auto:32768` reads 21% faster here | `--prefill` | 58,342 tokens | the long prompt | | --- | --- | --- | | `auto` (8,192) | 9.5 s | 548,737 tokens in 128.0 s, 4,287 tok/s | | `auto:16384` | 8.4 s | 548,737 tokens in 112.6 s, 4,874 tok/s | | `auto:32768` | 8.0 s | 565,121 tokens in 108.6 s, 5,201 tok/s | Same expert slots, same free VRAM. The cost I saw: checkpoints land on chunk boundaries, so the long prompt resumed from 32,768 tokens instead of 49,152. Where the time goes (`STRATA_PREFILL_TIMING=1`, the long prompt, ms): | phase | `auto` (8,192), 548,732 tokens | `auto:32768`, 565,116 tokens | | --- | --- | --- | | GPU timeline | 141,014 | 111,759 | | qsa select | 38,223 (27.1%) | 37,272 (33.4%) | | dequant | 18,129 (12.9%) | 4,257 (3.8%) | | gemm gate/up | 15,673 | 10,202 | | gemm down | 8,269 | 5,850 | | kv stage | 4,929 | 1,322 | | wait copy | 1,702 | 423 | | gdn recurrence | 9,567 | 9,906 | | host staging | 19,946 | 4,896 | Most of the gain is the `dequant` phase (the gather of each expert into its group slot, once per chunk) and the KV staging. On the 58K prompt `qsa select` is 3 to 5% of the read; at 565K it is a third. ## 2. Above ~135K cells the reference top-k costs more than the block scores `qsa_select_bench` is only built for HIP in `CMakeLists.txt`. I built it for CUDA (`target_link_libraries(qsa_select_bench PRIVATE strata_kernels CUDA::cudart)`) and ran it with 256 queries and a 1,048,576-cell capacity (ms per call): | context | block scores (TF32) | top-k | | --- | --- | --- | | 32,768 | 0.087 | 0.063 | | 131,072 | 0.289 | 0.229 (0.168 with `STRATA_TOPK_ACTIVE_ANY=1`) | | 400,000 | 0.835 | 1.06 | | 598,000 | 1.256 | 1.38 | | 1,000,000 | 2.269 | 3.13 | The scores are linear in the context. The top-k is not, and it passes the scores from 400K on: `fit` is `TK_T * TK_PER` = 33,792 blocks on NVIDIA, so beyond ~135K cells `qsa_block_topk` takes `qsa_block_topk_ref` whatever the active bound is (`STRATA_TOPK_ACTIVE_ANY=1` changed nothing at 400K and 598K). That kernel reads each query's score row six times (four radix passes, the count, the write) with 256 threads per query. So at a 1M context roughly half of `qsa select` is the reference top-k. I did not write a replacement. ## What did not help - **More key tiles per CTA in `block_scores_tc_kernel`** (`TC_ITER` 16 or 64 instead of 4): at 16, +2% at 598K and +5% at 1M; at 64, -2% and +7%. Both lose 8 to 64% at 131K and below. Scores bitwise identical. - **The query tile stored as its TF32 hi and lo parts** in shared memory, split once per CTA instead of once per key tile: 12 to 39% slower, at every context. Scores bitwise identical. - **`STRATA_TOPK_ACTIVE_ANY=1` in the engine**, default chunk: 9.53 s and 128.0 s, the same as without. "Bitwise identical" is an FNV-1a over the scores of the tested scorer, added to the bench, equal to the unchanged kernel's at 2,048, 32,768, 131,072, 400,000, 598,000 and 1,000,000 cells. ## Two things the bench shows on unchanged v0.1.38 code - The FP64 gate prints `FAIL` at 2,048 and 131,072 cells on this card: the tensor-core scorer's error is 1.0e-06 of the score scale, right at the floor (`max(4 * err_old, 1e-6 * scale)`), and `PASS` at 32,768 and 598,000 with 9.8e-07. If the bench becomes a CUDA ctest, that floor will flip. - At 400,000 cells, 253 of 256 selections are identical between the warp scorer and the tensor-core one (256 of 256 at the other five contexts). That is the "near-tie can select differently" the kernel's comment describes. Raw logs are available if you want them.
本站相关内容
相关页面的快捷入口。