Pull requests / #38
#38 moe: continuous adaptive expert cache eviction with exponential decay
closed · @code-martin · 0 评论 · 在 GitHub 查看
描述
### Summary Implements continuous, non-blocking MoE expert cache adaptation between VRAM resident slots and host RAM during multi-turn generation. ### Motivation Static expert placement profiles calibrated on generic corpora often yield low hit rates (~25–35%) when operating in specialized domains or conversations in other languages. Continuous adaptation enables the cache to dynamically follow conversation context. ### Key Changes - **Usage Frequency Tracking**: - Tracks expert routing activation counters per layer during token generation. - **Exponential Decay Smoothing**: - Applies a gentle 0.985 decay factor per round instead of abrupt reset (0.7), maintaining stable long-term resident profiles while remaining responsive to topic shifts. - **Hysteresis & Thrashing Prevention**: - Implements candidate and victim thresholds (gain +1.0) to avoid ping-pong thrashing between VRAM and RAM. - **Engine & Server Integration**: - Adds the ADAPT protocol command to generate.cpp. - Automatically triggers non-blocking adaptation rounds in serve/server.py at the end of each request. ### Testing - Tested on multi-turn conversations in other languages: observed over 5,600 successful expert swaps, with cache hit rate climbing from 25.8% to 63.0%.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。