Pull requests / #1674
#1674 fix: preserve pipeline token equality on 3x P4 (23 -> 28 tok/s)
open · @CC-David-CC · 0 comentários · No GitHub
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quants
Descrição
## 3x Tesla P4: 23.0 -> 28.6 tok/s with exact output At 4,096 prompt tokens, prose decode improved from **23.000485 to 28.606868 tok/s (+24.375063%)** with two-window pipelining. This is the same-day serial-versus-pipeline comparison on #1656 plus this guard; the pipeline speedup comes from #1656. ## Minimal correctness fix Add `!always_publish_` to the one-token self-commit condition in `Verifier::record_window()` and its matching shortcut in `Verifier::commit()`. Pipelined verifiers must leave recurrent state updates to the host verdict/snapshot and explicit commit path. Otherwise a one-token window can advance state during verification and then advance it again when the pipeline commits it. The complete code diff is **two guards and one explanatory comment in `src/core/verify.cpp`**. Ordinary serial verifiers retain their existing shortcut. No scheduler, kernel, configuration default, or benchmark files change. Found while testing [#1656](https://github.com/Niko1221/Strata/pull/1656). The mismatch also reproduced on its exact parent, `fb58e0dbc8399662c0e47c76578c6e878b14f6cf`, with a two-stage split, so it predates that PR. The guard is extracted from the earlier three-P4 work in [#1154](https://github.com/Niko1221/Strata/pull/1154), specifically [dd7045d](https://github.com/Niko1221/Strata/commit/dd7045d4b6a5143b73477b56c21b58bf3d6fa140). ## Measurements Measured 2026-10-09 on 3x Tesla P4, dual Xeon E5-2697 v3, 251.77 GiB usable RAM, CUDA 12.0, driver 580.178.04. Model: Qwen3.8-Flash-Next-GSQ-RCO **IQ3_XXS**, int8 KV, split `19,37`, 40,960-token context limit, 32,768 resident KV, 27 CPU workers, greedy decoding, MTP `--spec 2 --spec-min-p 0.5`, suffix drafting off. Static expert placement: `STRATA_IQ_MT_MIN=1 --pcie-frac 0 --adapt-every 0`. Both arms use the same corrected binary and expert placement, switching `pw=0` / `pw=2` per request. Three measured repetitions per arm; median engine decode throughput below. One initial serial/pipeline code warmup pair excluded. Order reversed for the middle repetition; zero reused prompt tokens throughout. Model loading excluded. | Workload | Prompt tokens | Generated tokens | Serial tok/s | Pipeline tok/s | Gain | | --- | ---: | ---: | ---: | ---: | ---: | | Prose | 4,096 | 256 | 23.000485 | 28.606868 | +24.375063% | | Prose | 512 | 256 | 23.377928 | 29.103477 | +24.491258% | | Code | 512 | 125 (EOS) | 25.461878 | 43.067806 | +69.146224% | | Counting | 512 | 256 | 26.400734 | 46.913942 | +77.699384% | 4K prose median whole-request latency: **33.755897 -> 31.510369 seconds**, a 6.652252% reduction. Decode throughput uses generated tokens divided by engine decode time, not whole-request time. A separate results-only PR will carry all per-run values, ranges, prompts, and reproduction files. ## Validation and scope - CUDA Release `sm_61` build succeeded on #1656 head `06abdfa50aef6e5068869bc7552fafa5e99b7b9d` plus exactly this patch. - All **32 fixed-build requests** matched their serial token reference: 2 warmups, 24 measured requests, and 6 forced-rollback/no-guess controls. Every 512-token-prompt result also matched the original unmodified serial output. - Forced rollback exercised 62 / 157 / 128 rollbacks on code / prose / counting. No crashes or OOMs. - Before the guard, the repeated 512-token prose pipeline run diverged at output token index 95 in all three repetitions. Disabling one-token self-commit globally also restored equality in the diagnostic screen. - These measurements validate IQ3_XXS on this CUDA host. Q8, sampled decoding, images, suffix drafting, dynamic expert adaptation, longer contexts, and general answer quality were not tested in this run. - HIP and SYCL builds are pending; keep this as a draft under the repository's requirement to build every touched backend before review.
No site
Links install, modelos, releases.