Pull requests / #10
#10 serve: read short prompt parts through the decode windows (~1 s faster first token)
closed · merged 2026-09-26 · @Mirtraxxx · 0 comentários · No GitHub
BenchmarksServer & APINVIDIA / CUDAWindows
Descrição
**Builds on #8.** This branch is #8 plus one commit (9d9351b); only that commit is new here. I can rebase it once #8 is merged or changed. ## What it does The batched prompt path has a fixed cost of ~300 ms per run however few tokens it reads: it streams every non-resident expert the chunk routes to over PCIe. On top of that, refilling the 1,131 expert-cache slots it borrows takes ~180 ms. So "hi" waited ~1.4 s for its first token. With #8, every chat follow-up paid the same fixed cost for a message of a few dozen tokens plus the 6-token assistant header. Now each part of the prompt (up to the turn checkpoint, then the header) picks its own path: - A part of at most `--short-read N` tokens (default 64, `0` = off) goes through the verify windows, S tokens at a time, as decode reads them: `ver.run` → `ver.commit` → `mtp.prefill`, with misses on the CPU pool. That's ~16 ms a token. - Longer parts, and parts with picture rows, stay batched. So a long first message is read batched, and its header still goes through the windows. - The slots are lent just before a batched run and refilled before a window read, so windows always see the whole expert cache. If the prompt ends on a batched part, the refill happens at the end, as before. - `STRATA_CKPT_REREAD` keeps every read batched, so it still compares a checkpoint against a batched re-read. ## Measured RTX 3090, Ryzen 7 5700X, 64 GB DDR4, Windows 11, Swift 1.5 IQ2_XS, `--serve --expert-cache auto --spec 4 --adapt-swaps 0`. Old and new engine back to back, same flags. | | old | new | |---|---:|---:| | First token, "hi" (13 tokens) | 1.37 s | **0.35 s** | | First token, 22-56 token prompts | 1.25-1.53 s | **0.41-0.85 s** | | Follow-up turn in a 5K-token chat (13-34 new tokens) | 1.11-1.25 s | **0.30-0.53 s** | | First turn, 5K tokens: the header part | 340-358 ms | 131 ms | | Long prompt, 7,685 tokens | 362 tok/s | 359 tok/s | | Decode, 966 tokens over 3 prompts | 52.5 tok/s | 53.9 tok/s | Output: greedy answers match the batched path except at near-ties, which the two paths round differently. On one prompt the first token flipped, and I dumped the logits to check it: "Think" 23.09 vs "Imagine" 23.02 through the windows, 23.03 vs 22.90 batched. The windows are the same computation decode uses. Only tested on Windows / one machine, text prompts. Picture rows always take the batched path, so image prompts should be unaffected, but I haven't tested that. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
No site
Links install, modelos, releases.