Issues / #1431

#1431 Feature request: tool-call emission recovery for low-bit quants in agent workloads (Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS)

open · @mechanicss · 1 comentários · No GitHub

AMD / HIPNVIDIA / CUDAModels & quantsLinux

Descrição

# Feature request: tool-call emission recovery for low-bit quants in agent workloads (Qwen3.8-Flash-Next, GSQ-RCO IQ3_XXS)

> Prepared as a complete case: reproduction, root-cause analysis, cross-references, a working workaround, and concrete feature proposals. All data (request payloads, mangled samples, logs, and our wrapper module) is available on request.

## Environment

- Model: **Qwen3.8-Flash-Next**, custom **GSQ-RCO IQ3_XXS** GGUF (2 shards + PLE), the public checkpoint.
- Engine: **Strata 0.1.40** (Linux, own CUDA build, `--kv int8`, `--max-context 262144`, resident experts).
- Hardware: RTX 4070 12 GB + 64 GB DDR3-1867, Intel Xeon E5-2696 v3 (the engine used 16 of 36 threads)
- Client: a coding-agent harness (DeepSeek Harness) sending **39 function schemas (~12K tokens)** plus a ~6K-token system prompt; total prompt ~16.5K tokens per turn.

## The problem

In agent workloads the model's tool-call emission fails stochastically, and Strata currently has no recovery path, so the harness receives broken text instead of a call. Two symptoms:

**1. Mangled emission → lands in content.** The model writes a broken variant of the native format and the parser cannot map it:

```
<tool_call>
<skill>
<name>
engineer-method
</</skill>
>
</name>
</name>
</function>
</tool_call>
<tool_call>
<bash>
<command>
pwd; ls -la; git status 2>&1 | head -5
</</>
</parameter>
<=>
ls; git status
</parameter
```

(Note the shorthand the broken model actually writes: `<skill>` / `<name>` instead of `<function='skill'>`; mismatched closers; occasional retried parameter values after the closers.)

Result: `finish_reason: stop`, no `tool_calls`, the markup leaks as the assistant message. The agent cannot act.

**2. Silent end.** The model closes its thinking block and emits EOS with **no answer and no call** (finish `stop`, empty content). Stochastic; appears with large tool-schema context; matches the degradation family of llama.cpp issue #28805 ("qwen4exp … 1 token then EOS at long context — stochastic threshold that moves with model quant / KV quant / n_ctx").

**Measured failure rate:** with a 39-tool schema set, a usable turn (call executed or meaningful answer) happens in **~30–50% of runs** over 20+ controlled replays of identical requests. With 3 tools the emission was once perfect; with 8 it already degrades — the failure clearly scales with schema/context pressure, not with prompt length alone.

## What proves this is recoverable

The **same weights** served by llama.cpp (master, `--jinja`) parse tool calls **perfectly**: 8+ consecutive runs of the exact same request (39 schemas, 16.5K prompt) — correctly parsed parallel calls (`skill(engineer-method)` + `skill(gate-checking)`, `bash` with the right command), zero leaks. llama.cpp succeeds because of (a) its PEG/grammar machinery that constrains and tolerantly maps the emission, and (b) its content→tool_calls lift fallback (see llama.cpp issue #30078, closed 2026-10-07 — the same failure family and the same fix approach; the reporter's numbers, 30–50% stochastic failure, match ours exactly).

We also built a **wrapper-level rescuer** for Strata (a filter between the output parser and the consumers): it buffers content once a `<tool_call` opener appears, and at the end of the turn repairs each segment — tool name resolved from structural slots (the tag right after the opener / `<name>` value / `<function='name'>`), parameters extracted from `<parameter=NAME>` and shorthand `<NAME>` openers known from the tool's JSON schema, values right-trimmed of closer junk and coerced by schema type; a confidence gate (unique name, all required params valid, no unknown explicit parameters, max 2 calls, never on cancel). It lifts repaired calls as real `ToolCall` events. **It works** — live log: `1 call(s) repaired from mangled content` — but wrapper-level repair cannot reach inside the emission process, so the useful-turn rate tops out around 30–50%.

## Feature requests, in order of impact

**1. Tolerant tool-call recovery in the output parser (content → tool_calls lift).**
When a `<tool_call>`-shaped region fails the strict parse, attempt schema-aware lenient repair (openers-based, as described above) and lift the result into real calls; or expose the raw region so the server layer can do it. llama.cpp #30078's reporter shipped exactly this kind of lift and it resolved their production failures. This alone would make agent workloads usable at 3-bit quants.

**2. Grammar-constrained emission option for tool calls.**
An opt-in mode (config/request flag) that constrains decoding once a `<tool_call>` opener appears: the structural skeleton (`<function=` + the request's tool names + `<parameter=` + closers) is grammar-masked, while parameter values stay free text. llama.cpp's `grammar_lazy` + trigger approach is the reference. This is the hard guarantee for low-bit quants — the model physically cannot produce a malformed call.

**3. Anti-silent-end option.**
When tools are present and the thinking block closes with no answer and no call, optionally re-prompt/resume once with a forced call opening (the server already has the `forced_call`/tool_choice machinery — a wrapper-level forced resume exists in our fork and fires correctly; the remaining gap is that the model sometimes writes prose inside the forced `<function=` — a grammar would close that too).

## Why we believe this is worth it

Qwen3.8-Flash-Next is new and the community is actively running it quantized (see llama.cpp #30078 created 2026-10-07 and ollama #18817 for GSQ-RCO import reports). Agent harnesses are the main real workload for this class of model, and tool-call emission is currently the first thing that breaks at low bits. Strata already reproduces llama.cpp-beating decode speeds on this exact model — emission recovery is the missing piece that would make it a complete agent engine.

## Data we can provide immediately

- Exact replay payloads (full request JSON with 39 schemas) that reproduce failures deterministically enough for A/B;
- Frozen mangled-output samples (raw bytes) from production runs;
- Engine + server logs around failures;
- Our wrapper rescuer module (Python, ~200 lines, 17 unit tests + a 90-case mutation corpus) as a reference implementation of the repair heuristics — happy to port the logic into the engine's parser if you point us at the right spot;
- A/B measurements before/after any candidate fix.

---

No site

Links install, modelos, releases.