SyncAI.news, a Varaisys broadcasting
Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
YC

Yunkai Chai, Tong Zhu, Xiaoye Qu, Xuyang Hu, Guanjie Chen, Qipeng Guo, Yu Cheng

· 1 min read

ResearcharXiv cs.CL

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

arXiv:2610.11575v1 Announce Type: new Abstract: Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-$k$ baseline by $1.6$ points on average across downstream tasks.

Original source

This story was published by arXiv cs.CL and written by Yunkai Chai, Tong Zhu, Xiaoye Qu, Xuyang Hu, Guanjie Chen, Qipeng Guo, Yu Cheng. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News