Issues / #1005

#1005 CRLF line-break tokens (`\r\n`, ids 317/845/8301) get far too little probability; everything else matches llama.cpp

open · @brenoperucchi · 3 comentarios · En GitHub

Setup & installNVIDIA / CUDAModels & quantsSecurityDocumentationWindows

Descripción

#803 was closed after its author traced the perplexity gap to CRLF line endings in the test text. The bug they found along the way has no issue of its own, so here it is, with a reproduction on a native GSQ-RCO pack.

On the same text, the same context and the same positions, Strata gives Windows line-break tokens a much higher NLL than the plain `\n` tokens they replace. All other tokens score the same.

## Setup

- RTX 5090 32 GB, Ryzen 9 5950X (AVX2), 96 GB DDR4-3200, Windows 11
- Strata 0.1.39 release, Swift IQ3_XXS (GSQ-RCO), the production config plus `--short-read 8192 --prompt-cache 0 --adapt-swaps 0`, scored with `STRATA_LOGPOS` (column 3)
- Text: the repo's own `docs/*.md` and `tools/*.py`, 8 chunks of 2048 tokens, second half of each chunk scored (8,184 positions). The CRLF copy is the same text with every `\n` replaced by `\r\n`, so the two token sequences differ only in the line-break ids.

## Result

Mean NLL per line-break token, same positions:

| Token in the LF text | NLL | Token in the CRLF text | NLL |
|---|---|---|---|
| `\n` (198), 239 times | 0.539 | `\r\n` (317) | 2.793 |
| `\n\n` (271), 64 times | 0.598 | `\r\n\r\n` (845) | 3.253 |
| `\n\n\n` (1358), 2 times | 2.974 | `\r\n\r\n\r\n` (8301) | 8.922 |
| all other tokens | 1.970 | all other tokens | 1.950 |

Perplexity goes from 6.807 (LF) to 7.292 (CRLF), and all of the increase sits on those three tokens. This matches what the #803 author reported on IQ3_S (mean NLL about 3.25 on the CRLF tokens against 0.39 for `\n`), and they saw it on 0.1.32 and 0.1.39, with three different packs and with `--pcie-frac 0` and `1.0`.

For reference, on the LF text Strata matches llama.cpp b11053 on the same GGUF: 6.807 vs 6.823 at 2K, 6.000 vs 5.987 at 8K (4 chunks of 8192), and 1.380 vs 1.380 on a Portuguese text. So the problem is specific to these ids.

## Why it matters

On Windows, CRLF reaches the model easily: files read from disk, tool output, text written in text mode. Short exact answers still came out right in #803, but the model reads every Windows line break as surprising, and any perplexity measured on such text looks worse than it is. We haven't measured what it does to long outputs.

## Workaround

Normalizing `\r\n` to `\n` before tokenizing avoids it (the #803 author does it in `ChatTemplate.render()`, and we added it to our client). That changes what the model sees and leaves the cause in place.

## Question

Is there anything in how these ids reach the model (the token embedding, the PLE n-gram lookup or its hashing) that treats them differently from llama.cpp? The #803 author checked the PLE hashing against llama.cpp's code and didn't find a difference. I can build from source on this machine and run any diagnostic build or dump you suggest.

En el sitio

Enlaces a install, modelos, releases.