Issues / #1385

#1385 Intermittent 'int' object has no attribute 'ranks' in tokenizer

open · @mustafaerdinc · 1 commentaires · Sur GitHub

Server & APINVIDIA / CUDAModels & quants

Description

During an OpenCode coding session, Strata logged this exception 12 times while preparing requests. OpenCode showed Bad Gateway, retried, and eventually finished. The Strata service did not restart and its cgroup OOM counters stayed at zero.
```text
serve/server.py: prepare -> encode_prompt -> self.tok.encode
tools/strata_tokenizer.py: _encode_plain -> self._bpe(mapped)
tools/strata_tokenizer.py: _bpe
    r = self.ranks.get((parts[i], parts[i + 1]))
AttributeError: 'int' object has no attribute 'ranks'
```

### Environment
- Ubuntu 26.04, x86_64
- CPython 3.14.4, built Aug 20 2026, GCC 15.2.0
- regex 2026.9.10
- RTX 4090, Qwen3.8-Flash-Next IQ3_S, context limit 131072
- Original deployment: commit 6f32ec070f23ced9f50e704d854d775da52591ab, engine 0.1.39
- Update tested: v0.1.40.1, commit 82f46a8c8f475f001ad76d92f58f4a4f8ffb0253, engine 0.1.40
The same exception also occurred during the update test, in a separate Python client process calling the tokenizer to count a long synthetic prompt. It happened before that process sent the long-prompt HTTP request. The client did not load Strata's native engine or NVML bindings. The server was running in another process at the time.

### What we checked
- The old and new `_bpe` implementations are identical. The failing line is in the short-word scan. `_bpe` reads `self.HEAP_MIN` earlier and does not assign to `self`.
- Fresh CPU-only processes encoded the same synthetic prompt successfully on both commits: 41829 tokens. Another 100 suffix variants passed on each commit.
- A later full update test passed, including a 41864-token model request. This did not establish that the intermittent error was fixed.
- Separate fresh Python probes reported JIT available but disabled, and GIL enabled. We did not measure those states inside the original failing server process.
Synthetic workload
### Synthetic workload
This describes the workload used in the CPU retests. It is not a reliable reproducer: those retests passed.
From the Strata checkout, with the affected pack's tokenizer loaded as `tok`:
```python
unit = "record: 1234567890 abcdefghijklmnop\n" * 100
unit_tokens = len(tok.encode(unit))
filler = unit * (40000 // unit_tokens + 1)
halfway = len(filler) // 2
content = (
    "SENTINEL=CEDAR917\n"
    + filler[:halfway]
    + "\nSENTINEL=MAPLE428\n"
    + filler[halfway:]
    + "\nSENTINEL=PINE652\n"
)
print(len(tok.encode(content)))
for case in range(100):
    tok.encode(content + "\nCPU_CASE=" + str(case))
```
We found #268, but that reports long-word tokenization latency rather than this exception. CPython python/cpython#156319 reports frame corruption with Tier 2 and active monitoring; we have not established those conditions here.

Is this a known tokenizer/runtime issue? If not, which process-level diagnostics or isolated Python-version comparison would help narrow it down? We currently have no confirmed root cause.

Sur le site

Liens install, modèles, releases.