Pull requests / #235

#235 verify: STRATA_LOGPOS hook for per-position log-probabilities (items 12-13 of #84)

closed · @enkynakamura · 0 コメント · GitHub で見る

Server & APIMulti-GPUAMD / HIPNVIDIA / CUDAWindows

本文

Adds Verifier::window_logprobs and a call to it in read_windows, enabled only when STRATA_LOGPOS=<path> is set. Every token read through the verify windows gets one tab-separated line: pos target logprob top top_logprob hit extra_logprob target_logprob_without_extra, where row t is the head's distribution at pos0 + t and the target is the prompt's own token at pos0 + t + 1. STRATA_LOGPOS_EXTRA (default 248046, <|im_end|>) is reported separately and taken out of the last column, because on raw text in a user turn a chat model puts much of its mass on the end-of-turn token.

Checked:

Built on v0.1.29 (d6708a4), CUDA 13.3, RTX 5080, Windows. A 32k probe dump (2931 rows) from this branch is byte-identical (md5 3e97790a…) to the dumps from v0.1.26, v0.1.28 and v0.1.29 built with the earlier revision of the hook.
Earlier, on the revision in #84: rows are aligned (top-1 equals the target only at shift 0), no periodicity per window, and a cross-check against llama.cpp on the same token ids (same top token in 9 of 11 positions, close log-probabilities).
Cost when enabled: 10.4 ms per token read from the windows, measured on 631c3d8. When disabled: one null check per window.

Not checked: the HIP build (the hook only uses cudaMemcpy/cudaMemcpyDeviceToHost, as the existing STRATA_DBG_NAN block does), and the layer-split path (forwarding plus OnDevice; I have a single GPU).

関連リンク

インストール・モデル・リリースへの站内リンク。