Pull requests / #500
#500 perf(cpu-pool): eliminate host serialization barrier for intermediate activation quantization with adaptive threshold
closed · @praveshkhatana · 0 Kommentare · Auf GitHub
BenchmarksModels & quantsWindows
Beschreibung
### Summary
In multi-token speculative verification passes (`run_split_multi` and `run_split_multi_native`), between phase 5 (gate/up projections) and phase 6 (down projections), intermediate activations $ff[t]$ are quantized (`act_quant_any` / `native_quant_h` into $a2$ or $hq$).
Previously, this intermediate activation quantization was executed **sequentially on the host thread** in a nested loop while all worker threads sat idle in a barrier:
```cpp
for (int e = 0; e < nb; ++e)
for (int t = 0; t < mjobs_[e].nt; ++t)
native_quant_h(f, split_multi_[e].ff[t], split_multi_[e].hq[t]);
```
This PR removes the host serialization barrier and introduces pooled parallel quantization protected by an adaptive threshold:
1. **Mode 7 in `ExpertPool::drain`**: Executes intermediate activation quantization tasks in parallel across pool workers using lock-free task claiming.
2. **Adaptive Threshold (`quant_threshold_ = std::max(8, n_workers)`)**:
- For standard generation and small verification windows ($\le \text{threshold}$), quantization executes sequentially on the host thread with **zero worker wake / spin-barrier overhead** (matching upstream baseline).
- Under heavy speculative bursts ($> \text{threshold}$), quantization tasks dispatch across idle pool workers via `run_phase(7)`.
- Configurable at runtime via `STRATA_POOL_QUANT_THRESH` (e.g. set to `0` to force parallel).
---
### Scope & Micro-Benchmarking Clarification
Internal engine timers (`pool multi ... gate/up / quantize / down ms/round`, profiled via `STRATA_VERIFY_PROFILE=1`) show that sequential intermediate quantization takes ~0.007 to 0.13 ms per round (<0.5% of round time on modern AVX2/AVX-512 CPUs).
This PR is therefore designed as a **thread-pipeline hygiene improvement** to eliminate the host wait during speculative bursts, rather than a macro tok/s booster. It guarantees zero regression on low-fanout verification passes via the adaptive threshold while cleanly utilizing available worker threads during multi-token bursts.
---
### Numerical Parity & Correctness
All bitwise parity tests pass with zero discrepancies:
- `iq_multi_parity`: **PASSED**
- `mmvq_multi_parity`: **PASSED**
- `quantize_act_parity`: **PASSED**
- `s2_expert_grouped_parity`: **PASSED**
- `native_expert_parity`: **PASSED**
Bit-for-bit identical outputs: intermediate quantizations write to independent, disjoint thread buffers (`split_multi_[e].a2[t]` / `split_multi_[e].hq[t]`), ensuring zero race conditions.
Mehr auf der Site
Links zu Install, Modellen, Releases.