SyncAI.news, a Varaisys broadcasting
Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism
JH

Jinfan He, Yunzhuo Liu, Kai Zhang, Weidong Han, Key, Rayying

· 1 min read

ResearcharXiv cs.AI

Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism

arXiv:2609.32398v1 Announce Type: new Abstract: The scaling of LLMs increasingly relies on MoE architectures to decouple active computation from total parameter count. However, the efficacy of MoE is often constrained by expert collapse and representation redundancy, both leading to underutilization of model capacity. To address these challenges, this paper proposes Distributional Orthogonalization Loss (DO-loss), an auxiliary regularization that shifts the focus from static weight diversity to dynamic routing behavior. By representing each expert's token assignment history as a high-dimensional binary load signature, DO-loss penalizes signature overlap to prevent expert collapse while encouraging functional specialization. To align this algorithmic design with system efficiency, we further introduce the Replica Expert Mechanism (REM), which improves load balancing through a two-tiered strategy: adjusting replica expert placement at the global-batch level and performing real-time token dispatching at the micro-batch level. Empirical evaluations demonstrate that our method outperforms the evaluated routing algorithms on downstream tasks for both 4.8BA0.5B and 30BA3B MoE models, while maintaining comparable training efficiency.

Original source

This story was published by arXiv cs.AI and written by Jinfan He, Yunzhuo Liu, Kai Zhang, Weidong Han, Key, Rayying. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News