Pull requests / #1700
#1700 generate, mtp: near the context's end the spec window and its drafts stay inside it
open · @midagedev · 0 comentarios · En GitHub
Server & APIAMD / HIPNVIDIA / CUDAModels & quantsLinux
Descripción
## Summary `strata generate --spec S` fails near the end of the context. The admission checks `prompt + max_new <= max_context`, but the loop still opens a full window there. So a request that fits ends with "ran out of context" (rc 2) and prints nothing. Near the end, `MtpDrafter::draft` can also draft into cell `max_cells`, past the page table. ## What changed - `src/program/generate.cpp`: a window takes only the rows that fit, `min(spec, max_context - p)`. The lookup chain stays inside them. - `src/core/mtp.cpp`: a draft round caps its steps at `max_cells - (p + a)`. Serve calls the same drafter, so it gets the cap too. The cap binds only at the end of the context. ## Extra Notes - CUDA 13.3 on Linux, RTX A6000, Qwen3.8-Flash-Next UD-Q4_K_XL, `fb58e0db` against this branch: - 4094-token prompt, `--max-context 4096 --max-new 2 --spec 4`: before, rc 2 and no output. After, rc 0 and 2 tokens. - 4084-token prompt, `--max-new 12 --spec 5 --suffix-draft 0 --mtp`, with a probe print: before, a round drafts into cell 4096 (`max_cells` 4096), then rc 2. After, its last cell is 4095, and the run ends with rc 0 and 12 tokens. - 512-token prompt, `--max-new 64 --spec 4 --mtp`, CPU share off: the same 64 tokens from both builds, twice each. - Not built: HIP and SYCL. The SYCL copies have the same lines. I can port the fix there. - I used Claude Code to find this and draft the change. I ran the tests myself.
En el sitio
Enlaces a install, modelos, releases.