Pull requests / #1676

#1676 verify: initialize remote expert helper state in both batch paths

open · @saikiran-rs · 0 comentarios · En GitHub

BenchmarksSetup & installAMD / HIPNVIDIA / CUDAWindowsLinux

Descripción

## Summary

When concurrent requests use `--remote-expert-opt`, batch verification can return
repeated characters or unrelated output. The batch graphs consume the helper's
mask and weighted sum, but `run_slot_rows` and `batch_poll` do not initialize the
helper with the current layer's routing weights or finish its work around the
CPU/helper pool call.

## What changed

Add guarded `remote_opt_->begin(h_w_, 0, S)` and `remote_opt_->end()` calls around
both batch pool calls, matching the existing solo verifier lifecycle. This is six
added lines in `src/core/verify.cpp`. The helper-disabled path executes no new
calls. Model weights, sampler, KV layout and GPU kernels are unchanged.

## Validation

Historical Linux HIP validation on v0.1.41 (`fb58e0d`) with RX 7900 XTX + RX 6800 XT,
Xeon E5-2696 v4, 128 GB DDR4 and Qwen3.8 Flash-Next UD-Q4_K_XL:

- Original helper-enabled three-client control passed 1/3 semantic replies.
  Disabling QFUSE also passed 1/3; disabling the helper passed 12/12.
- With this fix and the helper enabled, 18/18 initial/follow-up replies passed
  correct identity, arithmetic, meaningful-language and cross-session checks,
  with three native slots busy in every round. Another 12 configured-sampling
  replies passed.
- Patched solo/three-client sampled quality was 181/181 HellaSwag and 172/174
  WinoGrande, each out of 200, with zero invalid outputs. These scores compare
  modes of the patched runtime, not a before/after quality gain.
- All three configured 400K request windows recovered their own facts without
  cross-session leakage, including the generated output and API reservation.
- The patch applies to the current main source; whitespace checks pass.

These are correctness results from the recorded 2026-10-08 stack, which also has
separate Python API compatibility changes. They are not a new isolated speed
benchmark or a claim of bitwise/universal model-quality equivalence.
This PR is a draft: fresh PR-branch CUDA/HIP build checks and CUDA hardware testing
remain pending. No new model/GPU runs were performed for publication because the
machine is running a separate experiment. The original serving setup is unchanged.

## Extra notes

Related #1209/#1524 change batch layout/scheduling and #1636/#1637 change batch
MTP/adaptation. This change only supplies the missing helper lifecycle at the
two existing pool calls. It does not include the helper allowance setting from
#1458 or the performance report from #1459.

En el sitio

Enlaces a install, modelos, releases.