Pull requests / #38

#38 moe: continuous adaptive expert cache eviction with exponential decay

closed · @code-martin · 0 Kommentare · Auf GitHub

Server & APIModels & quants

Beschreibung

### Summary
Implements continuous, non-blocking MoE expert cache adaptation between VRAM resident slots and host RAM during multi-turn generation.

### Motivation
Static expert placement profiles calibrated on generic corpora often yield low hit rates (~25–35%) when operating in specialized domains or conversations in other languages. Continuous adaptation enables the cache to dynamically follow conversation context.

### Key Changes
- **Usage Frequency Tracking**:
  - Tracks expert routing activation counters per layer during token generation.
- **Exponential Decay Smoothing**:
  - Applies a gentle 0.985 decay factor per round instead of abrupt reset (0.7), maintaining stable long-term resident profiles while remaining responsive to topic shifts.
- **Hysteresis & Thrashing Prevention**:
  - Implements candidate and victim thresholds (gain +1.0) to avoid ping-pong thrashing between VRAM and RAM.
- **Engine & Server Integration**:
  - Adds the ADAPT protocol command to generate.cpp.
  - Automatically triggers non-blocking adaptation rounds in serve/server.py at the end of each request.

### Testing
- Tested on multi-turn conversations in other languages: observed over 5,600 successful expert swaps, with cache hit rate climbing from 25.8% to 63.0%.

Mehr auf der Site

Links zu Install, Modellen, Releases.