Pull requests / #84
#84 EXPERIMENTAL Rope scaling: contexts past the trained 262K (none / linear / YaRN)
closed · @j-luwierski · 0 comentários · No GitHub
Setup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Descrição
## What `--rope-scaling none|linear|yarn` - llama.cpp's types and flag names - extends the context past the model's trained 262,144 positions by rescaling the rotary angles: Position Interpolation for `linear`, ggml's `rope_yarn` (the corr_dims ramp and the mscale magnitude correction) for `yarn`. Engine 0.1.21; setup grows `--rope-scaling`/`--rope-scale` and 383K/512K context choices that pick the factor themselves. ## How it fits the engine - **The default path needs no kernel change.** Scaling is baked into the float64 host cos/sin table (`build_rope_table`), so `rope_neox_apply` and the indexer's pooled-key rotation cannot tell a scaled table from an unscaled one. - **The analytic paths** (`--native-rope`, the native indexer, prefill's `rope`) take the resolved constants (`freq_scale`, `corr_dims`, `ext_factor`, `attn_factor`) as kernel arguments. They are process constants set once before `session_init`, so CUDA-graph capture bakes them correctly; positions keep riding device memory (`st.pos_dev`). - One float32 helper (`rope_scaled_angle`), transcribed line-for-line from the vendored ggml with the MIT headers kept, serves all three analytic sites; one `RopeScaling` struct with a `rope_scaling_set()` startup setter follows the house pattern (`native_rope_set_enabled`, `mrope_table_set`). - Precedence: CLI over the model file's `qwen4exp.rope.*` keys over the struct defaults (the artifact ships no rope keys today, so that channel is a no-op for now). The scaling is fixed per process - K sits in the cache post-RoPE, so one cache must never mix two scalings; there is no per-request form. - `rope_parity` grows to 13 checks; docs (`DETAILS.md`) gain the "Context extension" section and a troubleshooting row. ## Validation - **Parity.** `none` is bit-identical to the old table builder; linear(4) is bit-exact against a float64 reference and row-equal to `none` at p/4; YaRN against a float64 transcription of ggml's spec; structural checks a tolerance cannot fake (`cos_tab[0] == mscale`; first pair extrapolates, last interpolates); an observability assertion; and native-vs-table agreement on the same input - the fixture that caught a real double-mscale bug in a first draft. - **No regression.** With no scaling flags the engine produces token-identical 64-token greedy output to the plain 0.1.20 build; the full parity suite and `serve.test_server` stay green. - **Recall past the trained end** (`tools/needle_bench.py`, depth 50%, temperature 0, RTX 4070 Ti SUPER + IQ3_XXS): 8/8 probes found - baseline (none) at 1k/32k/128K, then linear 2x and yarn 2x at 262k and 512k probe lengths; **the 512k probes are 527k prompt tokens, 2.01x the trained 262,144, recalled at mid-depth** (a 262k yarn probe re-ran green after the rebase onto 0.1.20). The raw runs stay on the development machine, out of this tree. ## Notes - Inside the trained range a scaled run is a slightly different model (YaRN's mscale applies at every position) - documented, with a startup warning when the context is within the trained range. - Pictures read the same scaled table; expected to compose, unmeasured (the recall runs are text).
No site
Links install, modelos, releases.