Pull requests / #1131
#1131 generate: the prompt loan keeps its residency bookkeeping without the token graph (#796 part B)
open · @j-luwierski · 0 Kommentare · Auf GitHub
BenchmarksSetup & installServer & APIMulti-GPUNVIDIA / CUDAModels & quants
Beschreibung
Part B of the #796 follow-up (part A landed in 0.1.40 as the one-allocation owned ring with exact pricing; part C follows separately as an advisory-only planner). The bug. The residency table (host_res/d_res) was built only inside the token-graph setup, so the prompt loan — whose lend marks the lent slots' experts not-resident and whose refill restores them — lost its bookkeeping under --no-token-graph or --no-capture: borrowing silently fell back to owned buffers the cache sizing had not reserved room for. Measured on the RTX 4070 Ti SUPER (16 GB), IQ3_XXS native pack, 12K prompt, --prefill auto: v0.1.40 under --no-token-graph dies at Prefill::init with "device buffers for a chunk of 8192 tokens do not fit (11 of 15938 MiB free)". The fix. The residency table is the loan's bookkeeping, so it is now built whenever the prompt path can borrow — an expert profile to fill the cache from, not --no-prefill-borrow, and a --prefill chunk — whether or not the token graph (its other reader) is captured. the decision is one pure predicate, strata::prefill::borrow_residency, with its mode matrix unit-tested; --no-prefill-borrow and chunkless runs still build nothing extra the token graph stays only one reader: the graph's hit wiring (the device hit count, the TokenHits fields, drive.d.host_res) remains behind the graph's own gate, so the decode pool's view is unchanged; the arena release/prefetch also runs in the borrowing-only case a staging failure keeps its old severity where the graph needs it and only falls back where the loan does (review fix): the table is freed, the loan scans see no residency state and the prompt path takes its own buffers, after one warning — prompt-path residency bookkeeping could not be staged (...); prefill borrowing is disabled for this run and the prompt path will use its own buffers. The graph's refusal (and the layer-split per-stage line) is byte-for-byte upstream's STRATA_TRACE names the loan (slots, first slot, experts marked non-resident) and the refill (restored, 0 left lent), so consecutive borrowed requests can be checked from the log No VRAM-planner policy is included — no plan, no startup enforcement, no accepted-plan contract. The pre-existing genuine errors (the residency-table staging refusal under the graph, the --spec device-table requirement, the serve start checks) are untouched. Measured (RTX 4070 Ti SUPER 16 GB, IQ3_XXS native pack, --spec 4, MTP): normal captured path unchanged (8192-token chunk, 2487–2489 borrowed slots, ~2199 tok/s prefill); --no-token-graph now borrows the same 2489 slots at ~2196 tok/s where 0.1.40 silently went owned and died at init; serve, two 8192-token requests back to back — each loan marks 2489 non-resident, each refill restores 2489 with 0 left lent; --no-prefill-borrow and the no-profile runs still own their buffers and build no residency state. Tests: borrow_residency_test (new, in ctest: the mode matrix + the staging-failure policy); full ctest 98% — the only failures are this machine's pre-existing ones (ple_parity missing fixture, expert_multi_test no AVX-512), identical on pristine main. Not covered here: multi-GPU (untouched; the split-stage fallback is code-verified), and --no-capture + borrowing on a native pack cannot start (pre-existing --spec requirement) — code path identical to the verified --no-token-graph case.
Mehr auf der Site
Links zu Install, Modellen, Releases.