Pull requests / #1676
#1676 verify: initialize remote expert helper state in both batch paths
open · @saikiran-rs · 0 comentarios · En GitHub
BenchmarksSetup & installAMD / HIPNVIDIA / CUDAWindowsLinux
Descripción
## Summary When concurrent requests use `--remote-expert-opt`, batch verification can return repeated characters or unrelated output. The batch graphs consume the helper's mask and weighted sum, but `run_slot_rows` and `batch_poll` do not initialize the helper with the current layer's routing weights or finish its work around the CPU/helper pool call. ## What changed Add guarded `remote_opt_->begin(h_w_, 0, S)` and `remote_opt_->end()` calls around both batch pool calls, matching the existing solo verifier lifecycle. This is six added lines in `src/core/verify.cpp`. The helper-disabled path executes no new calls. Model weights, sampler, KV layout and GPU kernels are unchanged. ## Validation Historical Linux HIP validation on v0.1.41 (`fb58e0d`) with RX 7900 XTX + RX 6800 XT, Xeon E5-2696 v4, 128 GB DDR4 and Qwen3.8 Flash-Next UD-Q4_K_XL: - Original helper-enabled three-client control passed 1/3 semantic replies. Disabling QFUSE also passed 1/3; disabling the helper passed 12/12. - With this fix and the helper enabled, 18/18 initial/follow-up replies passed correct identity, arithmetic, meaningful-language and cross-session checks, with three native slots busy in every round. Another 12 configured-sampling replies passed. - Patched solo/three-client sampled quality was 181/181 HellaSwag and 172/174 WinoGrande, each out of 200, with zero invalid outputs. These scores compare modes of the patched runtime, not a before/after quality gain. - All three configured 400K request windows recovered their own facts without cross-session leakage, including the generated output and API reservation. - The patch applies to the current main source; whitespace checks pass. These are correctness results from the recorded 2026-10-08 stack, which also has separate Python API compatibility changes. They are not a new isolated speed benchmark or a claim of bitwise/universal model-quality equivalence. This PR is a draft: fresh PR-branch CUDA/HIP build checks and CUDA hardware testing remain pending. No new model/GPU runs were performed for publication because the machine is running a separate experiment. The original serving setup is unchanged. ## Extra notes Related #1209/#1524 change batch layout/scheduling and #1636/#1637 change batch MTP/adaptation. This change only supplies the missing helper lifecycle at the two existing pool calls. It does not include the helper allowance setting from #1458 or the performance report from #1459.
En el sitio
Enlaces a install, modelos, releases.