Issues / #1703

#1703 Strata starts to digest a long prompt from the beginning

open · @fukc-gihtub · 1 commentaires · Sur GitHub

BenchmarksModels & quants

Description

It only happens occasionally, and I'm not sure whether it's an issue with the harness (OpenCode) or Strata. Only 51% of the 512k context is being used right now, and there's no compaction, but Strata suddenly "forgets" the prompt prefix and starts over:

```
[strata] writing a tool call: edit: 170 of max 32000 tokens, 84.9 tok/s, 3 s                                                        
[strata] writing a tool call: edit: 253 of max 32000 tokens, 83.3 tok/s, 4 s                                                        
[strata] writing a tool call: edit: 330 of max 32000 tokens, 81.2 tok/s, 5 s
[strata] writing a tool call: edit: 413 of max 32000 tokens, 81.2 tok/s, 6 s
[strata] done: 430 tokens in 6 s (81.8 tok/s) (stop, cancel=False), expert cache 95.3% hit (+7.2% of the routed experts over PCIe)
[strata] reading the prompt: 1,536 of 269,112 tokens, 1 s so far                                                                    
[strata] reading the prompt: 4,608 of 269,112 tokens, 3 s so far                                                                    
[strata] reading the prompt: 7,680 of 269,112 tokens, 5 s so far                                                                    
[strata] reading the prompt: 10,752 of 269,112 tokens, 7 s so far                                                                   
[strata] reading the prompt: 13,824 of 269,112 tokens, 9 s so far                                                                   
...
and so on.
```

Strata v0.1.41

Sur le site

Liens install, modèles, releases.