Pull requests / #1242
#1242 Isolate concurrent image positions and batch-MTP draft paths
open · @ilumn · 0 コメント · GitHub で見る
Setup & installNVIDIA / CUDAModels & quants
本文
With `--vision` and concurrent batching, admitting another image or a text request could overwrite the image-position table still used by an active slot's captured decode graph. This PR gives each slot its own table so another admission or slot reuse cannot change that slot's rotary positions. It applies the same isolation to batch-MTP drafting and rejects malformed vision records before they can affect a later valid request. This is the replacement for [#994](https://github.com/Niko1221/Strata/pull/994), which GitHub automatically closed after the upstream history rewrite. The recreation of the PR was requested [here](https://github.com/Niko1221/Strata/pull/994#issuecomment-6013987750). It ports original commit [10114f8](https://github.com/ilumn/Strata/commit/10114f8972c11cde0df5a90a94cd3a62b1f19916) onto the rewritten `main` at `82f46a8c8f475f001ad76d92f58f4a4f8ffb0253`. ### Changes - Give each batch slot its own image-position table on every stage. Copy the admitted request's positions before restoring draft state, and scope both batch-MTP drafter recording paths to that slot. - Guard verifier norm/RoPE and q/gate fusions when rows need slot-local tables; retain the fused global-position paths. Scope the per-token indexer fallback and grouped same-slot commits to the correct table. - Preserve the shared draft subset head's GGML type when binding a slot drafter. Talos testing found that the pointer was shared while its type stayed `-1`, so the first batch-MTP admission failed with `unsupported native MMVQ GGML type`. Copying that metadata fixes admission with the original subset head and the optional Q4 draft head. - Validate vision-record sizes, grids, and finite embedding values before use. Add parser regressions and CUDA graph checks for image and text/null tables, scope restoration, both visible devices, and a wrong-table negative control. ### Validation on Talos All **13 model-backed runs passed**, with each checked batch token stream matching its solo reference: | Configuration | Runs passed | Coverage | |---|---:|---| | One GPU, ordinary batching | 4 | Text batch, full text interleave, synthetic and encoder-produced mixed image/text | | One GPU, `--batch-mtp` | 5 | The same four cases, plus the optional Q4 draft-head batch check | | Two-GPU layer split | 4 | Text batch, supported staggered admission and next-turn comparisons, synthetic and encoder-produced mixed image/text | Mixed tests admit a different-grid image and text after the first image starts decoding. Greedy and seeded sampling, malformed-record rejection followed by valid recovery, cancellation, image-slot reuse, and image-to-text reuse all match their solo token references. Exactness controls include `STRATA_IQ_MT_MIN=1`, `--pcie-frac 0`, `--adapt-every 1000000`, and `--no-prefill-borrow`. - Full CUDA build for two RTX 5090s (`sm_120`) succeeded; the rebuild after the draft-head metadata correction also succeeded. - Final focused checks passed **3/3**: `vision_records_test`, `mrope_slot_graph_test`, and `rope_parity`, with both GPUs visible. - Strict parser compilation and ASan/UBSan passed (`detect_leaks=0`; leak checking was not run). - Full CUDA CTest inventory: **93 passed, 3 missing-fixture failures, 2 skips** (98 total). `ple_parity`, `expert_parity`, and `pool_test` lack their required development fixtures; `iq_parity_fixtures` and `iq_parity` skipped for unavailable fixtures. This inventory was run before the metadata correction; focused CUDA checks and affected model checks were rerun afterward. ### Scope and coverage limits - **Normal MTP works with a layer split**: the drafter lives on the last GPU, and the two-GPU solo references in these tests accepted MTP drafts. The existing engine guard disables only per-slot `--batch-mtp` when a layer split or helper GPU is configured; its current slot drafter setup uses the first stage's sessions and head. That guard is unchanged by this PR. - **Two-GPU `BYIELD`/resume was not validated.** In the upstream interleave harness, the idle split-stage prompt path bypassed the chunk-boundary yield handling. The supported staggered-admission and next-turn sections passed in a scoped derivative of that harness; the later yield/resume and checkpoint sections were not counted as passing on two GPUs. - Validation used 64-token decode caps and small test images; it does not establish all-model, long-context, or long-duration coverage. Historical GPU results from #994 are not counted as validation of this port.
関連リンク
インストール・モデル・リリースへの站内リンク。