Issues / #1168

#1168 Client-supplied stop sequences are ignored (OpenAI `stop` and Anthropic `stop_sequences`) — blocks agent/harness use

closed · @saadharis · 2 comentários · No GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quantsLinux

Descrição

## Summary

Strata ignores client-supplied **stop sequences** — both the OpenAI `stop` parameter on `/v1/chat/completions` and Anthropic `stop_sequences` on `/v1/messages`. Generation only ends on the model's EOS token ids, so any client relying on a stop string receives untruncated output.

## Environment

- Engine: **Strata v0.1.39**, source build on Linux (CUDA 12.6, sm_86)
- Model: `alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Strata-GGUF` → **IQ3_XXS** (75.8 GB)
- Hardware: 2× RTX 3090 (48 GB VRAM), 62 GB RAM — layer split CUDA0 layers 0-22 / CUDA1 layers 23-47
- Context: `262144`; endpoints exercised: `/v1/chat/completions`, `/v1/messages`

## Reproduction

**1. OpenAI `stop` is ignored**

```bash
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "strata",
  "messages": [{"role":"user","content":"List three colors, one per line."}],
  "max_tokens": 60, "temperature": 0, "reasoning_effort": "none",
  "stop": ["\n"]
}'
```

Expected: generation halts at the first newline → a single line.
Actual: `Red\nBlue\nGreen` — a full three-line answer, **byte-identical to the same request with `stop` omitted entirely**.

**2. The stop word is generated rather than honoured**

Same request with `"stop": ["STOP"]` → the literal text `STOP` appears in `content`, i.e. it was emitted as ordinary output.

**3. Anthropic `stop_sequences` is ignored too**

```bash
curl -s http://127.0.0.1:8080/v1/messages -H 'Content-Type: application/json' -d '{
  "model": "strata", "max_tokens": 60,
  "messages": [{"role":"user","content":"List three colors, one per line."}],
  "stop_sequences": ["\n"]
}'
```

Actual: the same three-line answer, and the response reports `"stop_sequence": null`.

**4. There is no hidden switch for it**

- No request-level stop read of any spelling in `serve/` (checked `stop`, `stop_sequences`, `.get("stop")`, `["stop"]`)
- `stop_tokens` in the source is `repeat_stop_tokens` — a repetition **loop-breaker**, not a user-facing stop
- No engine `--stop*` flag in `--help`, and no config key for it

## Impact

Clients that depend on stop sequences receive text they did not ask for:

- **lm-evaluation-harness `humaneval` scores 0.000** on this engine. The task sends `stop: ["\ndef", "\nclass", "\nif", "\nprint", ...]`. With the stop ignored, the model emits correct code and then **repeats the prompt inside a markdown fence**, which the completion extractor cannot parse. The same engine, same endpoint, same effort scores **0.874 on MMLU stem**, so this is not a capability limit.
- Code-completion clients, agent frameworks doing tool-call post-processing, and any harness that trims on a sentinel misbehave in the same way.

## Suggested fix

Truncate the generation stream on the first match of any supplied stop string (check the full detokenized buffer, not only token boundaries), end generation, and report `finish_reason: "stop"`. Wire `stop` on `/v1/chat/completions` and map Anthropic `stop_sequences` onto the same path, surfacing the matched string in the Anthropic response's `stop_sequence` field.

## Everything else has been excellent

For context, since this is the only blocker I've hit: ~99% of routed expert mass resident in VRAM on 2×3090, **105 tok/s decode and 2,383 tok/s prefill at 228K prompt tokens**, needle retrieval **3/3 at 10/50/90% depth out to 512K context**, and the mmproj vision tower verified end-to-end (3/3 on a synthetic ground-truth image). The Anthropic + Responses API surface on top is a real differentiator. Happy to test a build with this fix.

No site

Links install, modelos, releases.