SyncAI.news, a Varaisys broadcasting
MoEless: Efficient MoE LLM Serving with Serverless Experts
HY

Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang

· 1 min read

ResearcharXiv cs.AI

MoEless: Efficient MoE LLM Serving with Serverless Experts

arXiv:2603.06350v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

Original source

This story was published by arXiv cs.AI and written by Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News