Pull requests / #1455
#1455 serve: reuse prompt tokens across safe shared boundaries
open · @W1nge · 0 comentarios · En GitHub
Server & APINVIDIA / CUDAModels & quants
Descripción
Repeated and continued chat prompts currently run BPE over the entire rendered history. This adapts the shared-special-token-boundary reuse approach from gputier's #567 to current main; credit for the core algorithm belongs to that PR. Reuse is local to a service, bounded to four prompts (256K characters / 64K tokens each), and returns fresh token lists before image expansion. Requests containing protected literal spans keep the existing full encoder, preserving #537/#931 semantics. Prompts without a safe boundary and unsupported tokenizers also use full encoding. STRATA_PROMPT_REUSE=0 disables reuse. Validation: `python -m unittest serve.test_prompt_encoder serve.test_server -q` passed all 273 tests, covering overlapping specials, random edits, concurrent clients, tools, images, literal spans and actual prefix-BPE avoidance. No end-to-end latency claim is made for this port; full-model validation on current main is still outstanding, so this is a draft. Independent of the PLE and embedding-storage PRs; based on d5ea713 (v0.1.40.3). Validation update (2026-10-08): a complete current-main integration build including #1451, #1454, #1455, #1457 and #1465 passed four real IQ3_XXS model baseline/integration pairs on RTX 2080 Ti: default, GR_UNFUSED, PREFILL_BF16X2 and legacy RING_BYTES modes. Each used INT8 KV, 194 input tokens, two prefill chunks and 8 output tokens; all runs exited successfully and output IDs matched within each pair. A 256K configured-context startup and short generation also passed (this did not fill a 256K prompt). These are combined regression checks, not an isolated performance result or exhaustive state equality. Other hardware/backends remain untested. The earlier draft-only status reflected the absence of any current-main model run; now requesting review with these limits explicit.
En el sitio
Enlaces a install, modelos, releases.