Pull requests / #280

#280 RoPE: every kernel reads the session's float64 angle table

closed · @sergqwer · 0 评论 · 在 GitHub 查看

NVIDIA / CUDAWindows

描述

## Summary

Three RoPE implementations disagreed:

- the native decode/verify rope and the indexer's pooled keys computed `pos * powf(...)` with `cosf`/`sinf` under
  `--use_fast_math`: 0.0014 rad off at 32K, ~0.02 at 262K;
- the prompt path used precise float libm;
- the float64 angle table the session builds was read only by the non-native path.

`rope_table_set` now registers that table per device (`mrope.hpp`), and every kernel reads it through `rope_cs`, or
computes in float64 past its end. A position then has the same angle on every path.

## Measured

Measured on current main (0.1.29, d6708a4) with ISTA-DASLab's IQ2_XS, RTX 5090 32 GB, Ryzen 9 9950X3D, 128 GB DDR5, Windows 11. KL is first-token KL divergence (`STRATA_DUMP_FIRST_LOGITS`, #276); two runs of main itself differ by 0 on 5 of 6 short prompts and by 0.00013 on the sixth.

| prompt | KL, old angles vs the table | top-1 |
| --- | ---: | --- |
| 32K tokens, IQ2_XS | 0.0045 | same |
| 32K and 125K, NVFP4 pack (fork) | 0.0063 | same |

That is ~35x the run-to-run noise. Prompt and decode speed are unchanged (fork measurement).

## Switch

`STRATA_ROPE_LEGACY=1` brings the old float angles back.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01VZy1yKaDDiA8a7svdwaHio

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。