Issues / #1369
#1369 Two 32K prompts on one RTX 5090: overlapping them matched back-to-back wall time
open · @steve8697 · 0 comments · View on GitHub
BenchmarksNVIDIA / CUDAModels & quantsWindows
Description
## What this is One pair of measurements on a single RTX 5090, asking whether the default interleave (`STRATA_BATCH_DECODE_SHARE` 0.5) is the better schedule when **both** prompts are long. This is not a stall and not an admission crash. The engine stayed up. One run only, so the wall-clock comparison is a single sample. The log lines below are what that run actually printed. Related, and a different failure: #1118 was a 60-second stall on a DGX Spark while a long prompt was read beside a decoding slot. This machine did not stall. It finished, and the overlap did not beat running the two requests one after the other. ## Machine - Windows 11, RTX 5090 32 GB, driver 617.14, Ryzen 7 9700X. - Engine **0.1.40.2**, source `e8ca9af`. IQ3_S GSQ-RCO, KV int8, context 262144, vision on, no image sent. - `"parallel": 2`, `--batch-mtp`, `--spec 4`, `--prefill auto`, `--expert-cache auto` (10742 slots, 20.39 GiB at ready), `STRATA_PF_FUSED=1`. - `STRATA_BATCH_DECODE_SHARE` left at the default. The server was idle before each call. The only client was this test. ## Method Synthetic code plus an instruction to list observations. `temperature` 0, `reasoning_effort` none, `tool_choice` none, `max_tokens` 128. Each request had its own nonce at the start of the user text, so the fresh reads report `reused 0`. The local tokenizer aimed at 32768 prompt tokens. The engine counted **32526** on every request. The numbers below use the engine count. A slot handoff prints `reused` equal to the prompt just read and a near-zero `prompt_ms`. That line is the copy into the slot, not a second prefill, and it is not in the table. ## One request alone Client wall **6.6 s**. `finish_reason` `length`. ``` prompt 32526 tokens = 0 reused + 32526 read in 5663 ms (5743.2 tok/s), 128 generated in 886 ms (144.5 tok/s), drafts accepted 87 of 107 decode expert cache hit rate: 92.6% ``` Two of those back to back would be about **13.2 s**. ## Two requests at once Client walls **13.66 s** and **9.66 s**. Both `length`, 128 tokens. Makespan **13.7 s**, the same neighborhood as back to back. What the log shows: 1. The first prompt was read alone: `32526` tokens in **5478 ms** (5938 tok/s), then 1 token, then `slot 1 takes 32527 tokens (copied in 196.8 ms)`. 2. The second prompt was read in 3 parts while that slot decoded: `the prompt was read in 3 parts, the slots decoding 1203 ms between them`. The slot phase was `68 windows, avg 1.85 rows, 32.3 rows/s over 3899 ms`. `--batch-mtp` was on, and 1.85 rows across the active slots is barely more than one token per window. 3. The second prompt's own timing: `32526` tokens, `reused 0`, **6657 ms** (4886 tok/s). The extra ~1.2 s against the first prompt's 5.5 s matches the 1203 ms of inserted decode. 4. After both were in, `slot 0 gave back 32529 tokens ... in 180.9 ms`, and the remainder ran on the solo path: `124 generated in 688 ms (180.3 tok/s), drafts accepted 85 of 103`, hit rate 96.6%. So the first request spent its generation in a slot at about 32 rows/s while the second prompt was read, and the second request then decoded alone. The two long prefills did not run at the same time, and the two decodes did not share the GPU for long. ## Short prompts, for contrast Same machine, same server, earlier the same day. 44-token prompts, 160 generated tokens, one run. - Alone: client 1.74 s, decode 113.2 tok/s, drafts 81 of 127. - Two at once: clients 2.24 s and 2.28 s. Batch phase `97 windows, avg 3.27 rows, 199.4 rows/s over 1590 ms`. There, overlapping beat back to back. At 32K it did not. ## Question On a single 32 GB card, when the second request is also a long prompt, is the default half-chunk decode share still the schedule you want? A lower `STRATA_BATCH_DECODE_SHARE`, or reading the second long prompt with the slot paused, would move this pair closer to two full-speed prefills plus one fast solo decode. I have not tried other share values. Happy to rerun a share you would rather see measured.
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.