Pull requests / #1133
#1133 prefill: the advisory VRAM plan - the startup arithmetic says what it sees, and only warns (#796 part C)
open · @j-luwierski · 0 コメント · GitHub で見る
BenchmarksServer & APIMulti-GPUNVIDIA / CUDAModels & quantsSecurityWindowsLinux
本文
Part C of the #796 follow-up, independent of part B (clean off main 82f46a8). Part A's exact owned pricing (0.1.40) made the prompt path priceable before the expert cache is committed; this PR SAYS that arithmetic before the cache takes what is left of the VRAM — and warns instead of refusing. That is the one design change from the #796 draft Niko asked for: no startup refusal, no runtime contract, no second allocator. The engine's own sizing branches, its loan scan and its existing error paths stay the only authorities. What the planner says before the cache opens (single GPU): free VRAM at planning time, the user reserve, the draft head, --pipeline-windows; the prompt path's predicted mode and cost — owned buffers at the exact Prefill::bytes_needed_owned price beside what the sizing rule withholds (the 0.1.39 rule, or exact under STRATA_OWNED_PRICE=exact), or the loan the shared chunk/loan predictor (plan_lend_chunks, mirroring the runtime scan) derives; the predicted cache (the sizing walk, native sized slots included) and the #496 reserve adaptation the runtime is expected to make. Warnings, each with the knobs that make room (smaller --max-context, smaller --kv, --kv-resident where the model streams KV, smaller --prefill, smaller --vram-reserve-mib): a predicted shortfall, an over-large --expert-cache, a chunk no cache can lend, a zero-slot cache under --spec. After the cache opens: the free read is compared with the requirement (the auto sizing's own post-touch loop already wrote the slots), the prompt path is re-derived against the cache that actually exists, and a moved prediction — a loan that became owned buffers, a different auto chunk — is one warning line. At the loans: what the generate loan / the serve loan scan resolved is checked against the prediction; any difference is one warning: runtime prefill differs from the advisory VRAM plan: ... — a diagnostic; the run continues with the engine's own decision. Advisory stays pure (review fixes folded in): the planner prices at ITS OWN measured pinned share through the new share-pure Prefill pricing (bytes_needed_at_share, ring_slots_at_share, ring_default_slots_at, ring_cap_for_at — the same rules with the share in place of the global). It neither sets nor waits for Prefill::set_pinned_share; the runtime keeps its upstream call site below the cache sizing, and with the planner off the run reaches that site exactly as pristine main does. Env flags go through one parser (env_enabled): STRATA_ADVISORY_PLAN unset/any non-zero value enables, =0 silences; STRATA_POSTTOUCH (the deep WDDM check for explicit caches: one full cache write before the free read) stays off unless a non-zero value asks for it. Audited: no option mutation, no cache/chunk/ring/reserve/borrow change, no refusal anywhere in the advisory paths; the mirrored plan_lend_chunks is documented as diagnostic only — the runtime scan stays authoritative, and future drift surfaces as the mismatch warning. Loans keep their 0.1.39/0.1.40 layout and price; cache slot counts, chunks and reserves are unchanged. Measured — the no-regression proof (RTX 4070 Ti SUPER, IQ3_XXS native pack, 12K prompt, --prefill auto, --spec 4; three binaries: pristine main, this PR with STRATA_ADVISORY_PLAN=0, this PR with the planner on): pinned share < 0.9 (STRATA_ARENA_PIN_GIB=4, where a hoisted set_pinned_share would have shifted the runtime's ring pricing): all three binaries decision-identical — cache slots, chunk, loan, owned/borrowed; prefill 1410/1427/1417 tok/s (run noise) pinned share = 1.0 (default): chunk 8192 and loan 2489 slots / 3.99 GiB identical everywhere; cache slot counts jitter ±1% between ANY two runs on this desktop (a main-vs-main A/A spans the same spread); prefill 2194/2192/2191 tok/s the healthy run's prediction is exact: the 8192-token chunk, the 2489-slot loan, the opened 5539-slot / 8.98 GiB cache; a fixed --prefill 4096 predicts its 1559-slot loan exactly Tests: vram_plan_test (new, in ctest, GPU-less): the useful #796-draft coverage — the 13824/8704 bisection pins, exact sized-offset lending, the reserve as the user's knob, the re-derivations, the mismatch rules — with every !p.ok refusal assertion rewritten as has_warning() + usable numbers, plus the env parser. Full ctest 97%: the identical pre-existing failure set as pristine main. Not covered here: Windows/WDDM paging behavior (Linux/NVIDIA box; STRATA_POSTTOUCH exercised to its trace line only) and multi-GPU (the planner and every comparison run single-GPU only; a split's per-stage windows are booked by the engine as before).
関連リンク
インストール・モデル・リリースへの站内リンク。