Issues / #669

#669 RTX 5090, 1M context: --prefill auto:32768 reads a 598K prompt 21% faster, and above 135K cells the reference top-k costs more than the block scores

open · @gputier · 4 comments · View on GitHub

BenchmarksSetup & installServer & APIAMD / HIPNVIDIA / CUDAModels & quantsWindows

Description

Measurements of the prompt path at a 1M context on an RTX 5090, with what I tried and what did not help. No request attached: two observations you may want to act on, and three dead ends so that nobody repeats them.

## Setup

Windows 11, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB of RAM. Engine built from v0.1.38 (99f3dbd), Release, sm_120, CUDA 13.4, MSVC 14.44. Qwen3.8-Flash-Next GSQ-RCO IQ3_S, served by `serve/server.py` over `/v1/messages`.

Engine arguments: `--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --mtp <rt> --max-context 1048576 --kv int8 --kv-resident 32768 --rope-scaling yarn --rope-scale 4 --conversation-cache-mib 16384 --conversation-cache-slots 4`. The expert cache has 11,702 slots and 205 MiB of VRAM stay free, in every run below.

Two prompts, always in this order after a fresh start: 58,342 tokens read from 0, then 597,889 tokens whose first tokens are the first prompt (so it resumes from a checkpoint). Synthetic log lines, one question about a line in the middle; the answer was right in every run.

## 1. `--prefill auto:32768` reads 21% faster here

| `--prefill` | 58,342 tokens | the long prompt |
| --- | --- | --- |
| `auto` (8,192) | 9.5 s | 548,737 tokens in 128.0 s, 4,287 tok/s |
| `auto:16384` | 8.4 s | 548,737 tokens in 112.6 s, 4,874 tok/s |
| `auto:32768` | 8.0 s | 565,121 tokens in 108.6 s, 5,201 tok/s |

Same expert slots, same free VRAM. The cost I saw: checkpoints land on chunk boundaries, so the long prompt resumed from 32,768 tokens instead of 49,152.

Where the time goes (`STRATA_PREFILL_TIMING=1`, the long prompt, ms):

| phase | `auto` (8,192), 548,732 tokens | `auto:32768`, 565,116 tokens |
| --- | --- | --- |
| GPU timeline | 141,014 | 111,759 |
| qsa select | 38,223 (27.1%) | 37,272 (33.4%) |
| dequant | 18,129 (12.9%) | 4,257 (3.8%) |
| gemm gate/up | 15,673 | 10,202 |
| gemm down | 8,269 | 5,850 |
| kv stage | 4,929 | 1,322 |
| wait copy | 1,702 | 423 |
| gdn recurrence | 9,567 | 9,906 |
| host staging | 19,946 | 4,896 |

Most of the gain is the `dequant` phase (the gather of each expert into its group slot, once per chunk) and the KV staging. On the 58K prompt `qsa select` is 3 to 5% of the read; at 565K it is a third.

## 2. Above ~135K cells the reference top-k costs more than the block scores

`qsa_select_bench` is only built for HIP in `CMakeLists.txt`. I built it for CUDA (`target_link_libraries(qsa_select_bench PRIVATE strata_kernels CUDA::cudart)`) and ran it with 256 queries and a 1,048,576-cell capacity (ms per call):

| context | block scores (TF32) | top-k |
| --- | --- | --- |
| 32,768 | 0.087 | 0.063 |
| 131,072 | 0.289 | 0.229 (0.168 with `STRATA_TOPK_ACTIVE_ANY=1`) |
| 400,000 | 0.835 | 1.06 |
| 598,000 | 1.256 | 1.38 |
| 1,000,000 | 2.269 | 3.13 |

The scores are linear in the context. The top-k is not, and it passes the scores from 400K on: `fit` is `TK_T * TK_PER` = 33,792 blocks on NVIDIA, so beyond ~135K cells `qsa_block_topk` takes `qsa_block_topk_ref` whatever the active bound is (`STRATA_TOPK_ACTIVE_ANY=1` changed nothing at 400K and 598K). That kernel reads each query's score row six times (four radix passes, the count, the write) with 256 threads per query.

So at a 1M context roughly half of `qsa select` is the reference top-k. I did not write a replacement.

## What did not help

- **More key tiles per CTA in `block_scores_tc_kernel`** (`TC_ITER` 16 or 64 instead of 4): at 16, +2% at 598K and +5% at 1M; at 64, -2% and +7%. Both lose 8 to 64% at 131K and below. Scores bitwise identical.
- **The query tile stored as its TF32 hi and lo parts** in shared memory, split once per CTA instead of once per key tile: 12 to 39% slower, at every context. Scores bitwise identical.
- **`STRATA_TOPK_ACTIVE_ANY=1` in the engine**, default chunk: 9.53 s and 128.0 s, the same as without.

"Bitwise identical" is an FNV-1a over the scores of the tested scorer, added to the bench, equal to the unchanged kernel's at 2,048, 32,768, 131,072, 400,000, 598,000 and 1,000,000 cells.

## Two things the bench shows on unchanged v0.1.38 code

- The FP64 gate prints `FAIL` at 2,048 and 131,072 cells on this card: the tensor-core scorer's error is 1.0e-06 of the score scale, right at the floor (`max(4 * err_old, 1e-6 * scale)`), and `PASS` at 32,768 and 598,000 with 9.8e-07. If the bench becomes a CUDA ctest, that floor will flip.
- At 400,000 cells, 253 of 256 selections are identical between the warp scorer and the tensor-core one (256 of 256 at the other five contexts). That is the "near-tie can select differently" the kernel's comment describes.

Raw logs are available if you want them.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.