Issues / #669

#669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores

open · @gputier · 4 コメント · GitHub で見る

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows

本文

Measurements of the prompt path at a 1M context on an RTX 5090, with what I tried and what did not help. No request attached: two observations you may want to act on, and three dead ends so that nobody repeats them.

## Setup

Windows 11, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB of RAM. Engine built from v0.1.38 (99f3dbd), Release, sm_120, CUDA 13.4, MSVC 14.44. Qwen3.8-Flash-Next GSQ-RCO IQ3_S, served by `serve/server.py` over `/v1/messages`.

Engine arguments: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 1048576 --kv int8 --kv-resident 32768 --rope-scaling yarn --rope-scale 4 --conversation-cache-mib 16384 --conversation-cache-slots 4`. The expert cache has 11,702 slots and 205 MiB of VRAM stay free, in every run below.

Two prompts, always in this order after a fresh start: 58,342 tokens read from 0, then 597,889 tokens whose first tokens are the first prompt (so it resumes from a checkpoint). Synthetic log lines, one question about a line in the middle; the answer was right in every run.

## 1. `--prefill auto:32768` reads 21% faster here

| `--prefill` | 58,342 tokens | the long prompt |
| --- | --- | --- |
| `auto` (8,192) | 9.5 s | 548,737 tokens in 128.0 s, 4,287 tok/s |
| `auto:16384` | 8.4 s | 548,737 tokens in 112.6 s, 4,874 tok/s |
| `auto:32768` | 8.0 s | 565,121 tokens in 108.6 s, 5,201 tok/s |

Same expert slots, same free VRAM. The cost I saw: checkpoints land on chunk boundaries, so the long prompt resumed from 32,768 tokens instead of 49,152.

Where the time goes (`STRATA_PREFILL_TIMING=1`, the long prompt, ms):

| phase | `auto` (8,192), 548,732 tokens | `auto:32768`, 565,116 tokens |
| --- | --- | --- |
| GPU timeline | 141,014 | 111,759 |
| qsa select | 38,223 (27.1%) | 37,272 (33.4%) |
| dequant | 18,129 (12.9%) | 4,257 (3.8%) |
| gemm gate/up | 15,673 | 10,202 |
| gemm down | 8,269 | 5,850 |
| kv stage | 4,929 | 1,322 |
| wait copy | 1,702 | 423 |
| gdn recurrence | 9,567 | 9,906 |
| host staging | 19,946 | 4,896 |

Most of the gain is the `dequant` phase (the gather of each expert into its group slot, once per chunk) and the KV staging. On the 58K prompt `qsa select` is 3 to 5% of the read; at 565K it is a third.

## 2. Above ~135K cells the reference top-k costs more than the block scores

`qsa_select_bench` is only built for HIP in `CMakeLists.txt`. I built it for CUDA (`target_link_libraries(qsa_select_bench PRIVATE strata_kernels CUDA::cudart)`) and ran it with 256 queries and a 1,048,576-cell capacity (ms per call):

| context | block scores (TF32) | top-k |
| --- | --- | --- |
| 32,768 | 0.087 | 0.063 |
| 131,072 | 0.289 | 0.229 (0.168 with `STRATA_TOPK_ACTIVE_ANY=1`) |
| 400,000 | 0.835 | 1.06 |
| 598,000 | 1.256 | 1.38 |
| 1,000,000 | 2.269 | 3.13 |

The scores are linear in the context. The top-k is not, and it passes the scores from 400K on: `fit` is `TK_T * TK_PER` = 33,792 blocks on NVIDIA, so beyond ~135K cells `qsa_block_topk` takes `qsa_block_topk_ref` whatever the active bound is (`STRATA_TOPK_ACTIVE_ANY=1` changed nothing at 400K and 598K). That kernel reads each query's score row six times (four radix passes, the count, the write) with 256 threads per query.

So at a 1M context roughly half of `qsa select` is the reference top-k. I did not write a replacement.

## What did not help

- **More key tiles per CTA in `block_scores_tc_kernel`** (`TC_ITER` 16 or 64 instead of 4): at 16, +2% at 598K and +5% at 1M; at 64, -2% and +7%. Both lose 8 to 64% at 131K and below. Scores bitwise identical.
- **The query tile stored as its TF32 hi and lo parts** in shared memory, split once per CTA instead of once per key tile: 12 to 39% slower, at every context. Scores bitwise identical.
- **`STRATA_TOPK_ACTIVE_ANY=1` in the engine**, default chunk: 9.53 s and 128.0 s, the same as without.

"Bitwise identical" is an FNV-1a over the scores of the tested scorer, added to the bench, equal to the unchanged kernel's at 2,048, 32,768, 131,072, 400,000, 598,000 and 1,000,000 cells.

## Two things the bench shows on unchanged v0.1.38 code

- The FP64 gate prints `FAIL` at 2,048 and 131,072 cells on this card: the tensor-core scorer's error is 1.0e-06 of the score scale, right at the floor (`max(4 * err_old, 1e-6 * scale)`), and `PASS` at 32,768 and 598,000 with 9.8e-07. If the bench becomes a CUDA ctest, that floor will flip.
- At 400,000 cells, 253 of 256 selections are identical between the warp scorer and the tensor-core one (256 of 256 at the other five contexts). That is the "near-tie can select differently" the kernel's comment describes.

Raw logs are available if you want them.

関連リンク

インストール・モデル・リリースへの站内リンク。