Pull requests / #1419
#1419 bench: replay token IDs and verified windows in serve decode
open · @imanu86 · 0 コメント · GitHub で見る
BenchmarksServer & APINVIDIA / CUDAWindows
本文
This adds opt-in forced-token and verified-window replay controls to the serial and pipelined GEN/GENI serve loops. They are laboratory tools for comparing a fixed continuation when ordinary greedy runs diverge because of arithmetic or expert-placement differences. They do not assess generated-text quality. - `STRATA_FORCE_IDS=<file>[;<file>...]`: request k uses file k's token IDs, replacing target picks and limiting `max_new` to the list length. Files may contain little-endian int32 values or a decimal text list. Requests beyond the file list are not forced. - `STRATA_FORCE_TRACE=<file>`: append `W <request> <position> <rows> <accepted> <row tokens...>` for each verified window. - `STRATA_FORCE_WINDOWS=<file>`: replay matching window rows from a reference trace. A match needs the same position and first token and must fit the current window capacity. The drafter still runs. Speculative windows that do not match retain their normal path. - A per-request `strata work:` summary reports window widths, accepted/offered drafts, emitted tokens, replaced target picks on verified rows, and replay hits/misses. Scope and interpretation: The default path keeps normal token selection, but it does gain accounting and the stderr summary. This is not a claim of zero overhead. Equal forced output, window histograms and accepted counts help control the verified workload; they do not prove identical expert residency, cache adaptation, PLE stalls, speculative rollbacks, allocations or overlapping work. Timings still need matched conditions and repeated comparisons. Forced output cannot be used as a correctness or model-quality pass. These hooks cover the GEN/GENI serial and pipeline loops in this change; no concurrent batch-slot replay or standalone CLI replay coverage is claimed. Inputs are intended to be trusted, validated laboratory fixtures. The current parser has permissive fallback/skip behavior for missing or malformed data; stricter fail-closed fixture validation is a review item, not an implemented guarantee. Validation: - Branch `pr/forced-bench`, head `4780a0cd6bd33b944243c0862255a2f2c0393232`, three commits over `82f46a8c8f475f001ad76d92f58f4a4f8ffb0253`. `git diff --check 82f46a8..HEAD` passes. - The corresponding hooks have been used in the local fork on Windows with RTX 3060 + RTX 2080 Ti, split layers, MTP, two pipeline windows, actual prompts 131248/131239 tokens and 1024 forced output tokens. Historical logs reproduce verified-window histograms of `1:191 2:116 3:68 4:146` and `1:254 2:175 3:75 4:95 5:1`. These are evidence of exercised hooks in the fork, not fresh validation of this exact PR head. - No fresh build or GPU run of this branch was performed for publication. This branch still needs compilation and checks for unforced behavior, trace round-trip, mismatched traces and request indexing on current upstream. Suggested replay procedure: ```text # Record valid continuation IDs from an ordinary reference run. # Force that continuation and record a fresh trace file: STRATA_FORCE_IDS="ids_0.txt;ids_1.txt" STRATA_FORCE_TRACE=ref_trace.txt strata ... # Compare treatments using the same validated IDs, trace and requests: STRATA_FORCE_IDS="ids_0.txt;ids_1.txt" STRATA_FORCE_WINDOWS=ref_trace.txt strata ... ``` The trace is append-only: use a new path or intentionally cleared file when recording a reference. Inspect the work counters and the rest of the runtime state before interpreting a timing difference.
関連リンク
インストール・モデル・リリースへの站内リンク。