Issues / #684

#684 [Question] Prompt reProcessing

closed · @ukrolelo · 2 评论 · 在 GitHub 查看

BenchmarksModels & quants

描述

is it correctly working that's it reprocessing the prompts from 8192 almost every time? It's same session.
Question: how to decrease amount of reprocessing every time? No subagents,no subtasks Thx


Logs:
[strata] answering: 20465 of max 172610 tokens, 141.9 tok/s, 151 s
[strata] answering: 20613 of max 172610 tokens, 142.0 tok/s, 152 s
[strata] answering: 20771 of max 172610 tokens, 142.1 tok/s, 153 s
[strata] answering: 20920 of max 172610 tokens, 142.1 tok/s, 154 s
[strata] answering: 21074 of max 172610 tokens, 142.2 tok/s, 155 s
[strata] answering: 21225 of max 172610 tokens, 142.2 tok/s, 156 s
[strata] answering: 21363 of max 172610 tokens, 142.2 tok/s, 157 s
[strata] done: 21471 tokens in 158 s (142.2 tok/s) (stop, cancel=False), expert cache 99.3% hit
[strata] reading the prompt: 8,192 of 50,605 tokens, 2 s so far
[strata] reading the prompt: 16,384 of 50,605 tokens, 4 s so far
[strata] reading the prompt: 27,244 of 50,605 tokens, 7 s so far
[strata] reading the prompt: 35,436 of 50,605 tokens, 8 s so far
[strata] reading the prompt: 43,628 of 50,605 tokens, 10 s so far
[strata] reading the prompt: 50,600 of 50,605 tokens, 12 s so far
[strata] thinking: 79 of max 151387 tokens, 105.0 tok/s, 13 s
[strata] writing a tool call: skill_view: 186 of max 151387 tokens, 105.6 tok/s, 14 s
[strata] done: 203 tokens in 14 s (108.2 tok/s) (stop, cancel=False), expert cache 98.1% hit
[strata] reading the prompt: 58,792 of 61,790 tokens, 2 s so far
[strata] reading the prompt: 61,789 of 61,790 tokens, 4 s so far
[strata] done: 123 tokens in 5 s (130.8 tok/s) (stop, cancel=False), expert cache 98.5% hit
[strata] reading the prompt: 63,573 of 63,574 tokens, 1 s so far
[strata] thinking: 109 of max 138418 tokens, 110.7 tok/s, 2 s
[strata] thinking: 213 of max 138418 tokens, 107.0 tok/s, 3 s
[strata] thinking: 339 of max 138418 tokens, 113.1 tok/s, 4 s
[strata] thinking: 460 of max 138418 tokens, 114.8 tok/s, 5 s
[strata] thinking: 556 of max 138418 tokens, 110.9 tok/s, 6 s
[strata] writing a tool call: execute_code: 667 of max 138418 tokens, 110.8 tok/s, 7 s
[strata] writing a tool call: execute_code: 811 of max 138418 tokens, 115.4 tok/s, 8 s
[strata] done: 953 tokens in 9 s (118.9 tok/s) (stop, cancel=False), expert cache 99.0% hit
[strata] thinking: 58 of max 201339 tokens, 135.1 tok/s, 1 s
[strata] done: 155 tokens in 2 s (115.6 tok/s) (stop, cancel=False), expert cache 99.1% hit
[strata] thinking: 63 of max 201398 tokens, 125.5 tok/s, 1 s
[strata] done: 82 tokens in 1 s (120.8 tok/s) (stop, cancel=False), expert cache 99.5% hit
[strata] reading the prompt: 8,192 of 64,107 tokens, 2 s so far
[strata] reading the prompt: 16,384 of 64,107 tokens, 4 s so far
[strata] reading the prompt: 27,244 of 64,107 tokens, 6 s so far
[strata] reading the prompt: 35,436 of 64,107 tokens, 8 s so far
[strata] reading the prompt: 43,628 of 64,107 tokens, 10 s so far
[strata] reading the prompt: 51,820 of 64,107 tokens, 12 s so far
[strata] reading the prompt: 60,012 of 64,107 tokens, 14 s so far
[strata] reading the prompt: 64,102 of 64,107 tokens, 15 s so far
[strata] thinking: 76 of max 137885 tokens, 98.9 tok/s, 16 s

Config:
{
 "exe": "A:\\0_strata\\engine\\strata.exe",
 "args": [
  "--pack",
  "A:\\Strata-data\\packs\\iq3_s",
  "--native",
  "A:\\Strata-data\\models\\IQ3_S\\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf",
  "--ple-gguf",
  "A:\\Strata-data\\models\\IQ3_S\\Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf",
  "--expert-profile",
  "A:\\0_strata\\data\\expert-profile.bin",
  "--expert-cache",
  "auto",
  "--prefill",
  "auto",
  "--spec",
  "4",
  "--mtp",
  "A:\\Strata-data\\mtp\\rt",
  "--max-context",
  "202000",
  "--kv",
  "int8",
  "--kv-resident",
  "32768",
  "--vision",
  "--vram-reserve-mib",
  "700",
  "--pcie-frac",
  "0.55",
  "--spec-min-p",
  "0.70",
  "--pool-workers",
  "15"
 ],

RAM usage 73 GB/256 GB
VRAM usage 47 GB/48 GB

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。