Pull requests / #1692
#1692 bench: existing RoPE vs YaRN at 512K and 1M, with setup guide
open · @CC-David-CC · 0 comentários · No GitHub
BenchmarksSetup & installNVIDIA / CUDAModels & quantsDocumentationWindowsLinux
Descrição
 ## Summary Measures Strata's existing YaRN on untouched main `fb58e0db`; no inference changes or converter dependency. ISTA IQ3_XXS, FP16 KV, RTX PRO 6000 Blackwell. One fresh three-code retrieval prompt per condition: | Configuration | Correct pairs | Prefill | Decode | |---|---:|---:|---:| | Ordinary RoPE, 512K | 3/3 | 95.7 s | 103.6 tok/s | | Ordinary RoPE, 1M | 0/3 | 235.5 s | 82.4 tok/s | | YaRN 4x, 1M | 3/3 | 232.8 s | 82.7 tok/s | The 1M input hashes and executable match. Ordinary RoPE retained the names/numbers but swapped their associations; exact answers are shown in the figure. Replies were only 22 tokens. ## Use existing YaRN Linux: `./setup.sh --setup --context 1048576 --rope-scaling yarn --rope-scale 4` Windows: `START-HERE.bat --setup --context 1048576 --rope-scaling yarn --rope-scale 4` Stop the running server first. `--setup` saves these settings and starts the selected model; later launches use `./setup.sh` or `START-HERE.bat`. Includes engine flags, a 128-token API example, output headroom and cache compatibility guidance in the [setup and reproduction guide](https://github.com/CC-David-CC/Strata-a5500/blob/7fd0910a2fe2b2d735b3dce66416ad21b5b5dcd7/bench/results/2026-10-09-rope-yarn-long-context/README.md). ## Scope This is a small reproducible probe, not a validated 512K switching cutoff or general 1M quality guarantee. Retain normal defaults within the declared 262K context; beyond that, follow setup's existing experimental YaRN policy. Fresh YaRN is not cache migration. Adds raw measurements/hashes, the standalone probe, PNG/SVG figure and plot source. No weights, saved sessions, or credentials. Related measurements: #348 and #781.
No site
Links install, modelos, releases.