Issues / #1431
#1431 Feature request: tool-call emission recovery for low-bit quants in agent workloads (Qwen3.8-Flash-Next GSQ-RCO IQ3_XXS)
open · @mechanicss · 1 comments · View on GitHub
AMD / HIPNVIDIA / CUDAModels & quantsLinux
Description
# Feature request: tool-call emission recovery for low-bit quants in agent workloads (Qwen3.8-Flash-Next, GSQ-RCO IQ3_XXS)
> Prepared as a complete case: reproduction, root-cause analysis, cross-references, a working workaround, and concrete feature proposals. All data (request payloads, mangled samples, logs, and our wrapper module) is available on request.
## Environment
- Model: **Qwen3.8-Flash-Next**, custom **GSQ-RCO IQ3_XXS** GGUF (2 shards + PLE), the public checkpoint.
- Engine: **Strata 0.1.40** (Linux, own CUDA build, `--kv int8`, `--max-context 262144`, resident experts).
- Hardware: RTX 4070 12 GB + 64 GB DDR3-1867, Intel Xeon E5-2696 v3 (the engine used 16 of 36 threads)
- Client: a coding-agent harness (DeepSeek Harness) sending **39 function schemas (~12K tokens)** plus a ~6K-token system prompt; total prompt ~16.5K tokens per turn.
## The problem
In agent workloads the model's tool-call emission fails stochastically, and Strata currently has no recovery path, so the harness receives broken text instead of a call. Two symptoms:
**1. Mangled emission → lands in content.** The model writes a broken variant of the native format and the parser cannot map it:
```
<tool_call>
<skill>
<name>
engineer-method
</</skill>
>
</name>
</name>
</function>
</tool_call>
<tool_call>
<bash>
<command>
pwd; ls -la; git status 2>&1 | head -5
</</>
</parameter>
<=>
ls; git status
</parameter
```
(Note the shorthand the broken model actually writes: `<skill>` / `<name>` instead of `<function='skill'>`; mismatched closers; occasional retried parameter values after the closers.)
Result: `finish_reason: stop`, no `tool_calls`, the markup leaks as the assistant message. The agent cannot act.
**2. Silent end.** The model closes its thinking block and emits EOS with **no answer and no call** (finish `stop`, empty content). Stochastic; appears with large tool-schema context; matches the degradation family of llama.cpp issue #28805 ("qwen4exp … 1 token then EOS at long context — stochastic threshold that moves with model quant / KV quant / n_ctx").
**Measured failure rate:** with a 39-tool schema set, a usable turn (call executed or meaningful answer) happens in **~30–50% of runs** over 20+ controlled replays of identical requests. With 3 tools the emission was once perfect; with 8 it already degrades — the failure clearly scales with schema/context pressure, not with prompt length alone.
## What proves this is recoverable
The **same weights** served by llama.cpp (master, `--jinja`) parse tool calls **perfectly**: 8+ consecutive runs of the exact same request (39 schemas, 16.5K prompt) — correctly parsed parallel calls (`skill(engineer-method)` + `skill(gate-checking)`, `bash` with the right command), zero leaks. llama.cpp succeeds because of (a) its PEG/grammar machinery that constrains and tolerantly maps the emission, and (b) its content→tool_calls lift fallback (see llama.cpp issue #30078, closed 2026-10-07 — the same failure family and the same fix approach; the reporter's numbers, 30–50% stochastic failure, match ours exactly).
We also built a **wrapper-level rescuer** for Strata (a filter between the output parser and the consumers): it buffers content once a `<tool_call` opener appears, and at the end of the turn repairs each segment — tool name resolved from structural slots (the tag right after the opener / `<name>` value / `<function='name'>`), parameters extracted from `<parameter=NAME>` and shorthand `<NAME>` openers known from the tool's JSON schema, values right-trimmed of closer junk and coerced by schema type; a confidence gate (unique name, all required params valid, no unknown explicit parameters, max 2 calls, never on cancel). It lifts repaired calls as real `ToolCall` events. **It works** — live log: `1 call(s) repaired from mangled content` — but wrapper-level repair cannot reach inside the emission process, so the useful-turn rate tops out around 30–50%.
## Feature requests, in order of impact
**1. Tolerant tool-call recovery in the output parser (content → tool_calls lift).**
When a `<tool_call>`-shaped region fails the strict parse, attempt schema-aware lenient repair (openers-based, as described above) and lift the result into real calls; or expose the raw region so the server layer can do it. llama.cpp #30078's reporter shipped exactly this kind of lift and it resolved their production failures. This alone would make agent workloads usable at 3-bit quants.
**2. Grammar-constrained emission option for tool calls.**
An opt-in mode (config/request flag) that constrains decoding once a `<tool_call>` opener appears: the structural skeleton (`<function=` + the request's tool names + `<parameter=` + closers) is grammar-masked, while parameter values stay free text. llama.cpp's `grammar_lazy` + trigger approach is the reference. This is the hard guarantee for low-bit quants — the model physically cannot produce a malformed call.
**3. Anti-silent-end option.**
When tools are present and the thinking block closes with no answer and no call, optionally re-prompt/resume once with a forced call opening (the server already has the `forced_call`/tool_choice machinery — a wrapper-level forced resume exists in our fork and fires correctly; the remaining gap is that the model sometimes writes prose inside the forced `<function=` — a grammar would close that too).
## Why we believe this is worth it
Qwen3.8-Flash-Next is new and the community is actively running it quantized (see llama.cpp #30078 created 2026-10-07 and ollama #18817 for GSQ-RCO import reports). Agent harnesses are the main real workload for this class of model, and tool-call emission is currently the first thing that breaks at low bits. Strata already reproduces llama.cpp-beating decode speeds on this exact model — emission recovery is the missing piece that would make it a complete agent engine.
## Data we can provide immediately
- Exact replay payloads (full request JSON with 39 schemas) that reproduce failures deterministically enough for A/B;
- Frozen mangled-output samples (raw bytes) from production runs;
- Engine + server logs around failures;
- Our wrapper rescuer module (Python, ~200 lines, 17 unit tests + a 90-case mutation corpus) as a reference implementation of the repair heuristics — happy to port the logic into the engine's parser if you point us at the right spot;
- A/B measurements before/after any candidate fix.
---
Related on strata.com
Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.