
QY
Qingao Yi, Jiaang Duan, Jun Zhang, Haiyan Zhao, Shiyou Qian, Dingyu Yang, Jian Cao, Jinghua Tang
· 1 min read
ResearcharXiv cs.AI
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
arXiv:2511.10333v2 Announce Type: replace-cross
Abstract: Training large language models (LLMs) at scale incurs substantial communication overhead, while static gradient compression cannot adapt to gradient evolution and may degrade model quality. We propose EDGC, an entropy-driven dynamic gradient compression framework that adapts compression ranks to gradient entropy during training. EDGC combines efficient entropy estimation through gradient sampling, a theoretical model relating entropy to compression rank under a bounded-error constraint, and window-based rank adjustment across pipeline stages. Experiments on 32-V100 and 64-H100 GPU clusters training GPT2 models with 2.5B and 12.1B parameters show that EDGC reduces communication latency by up to 46.45% and end-to-end training time by 16.13%, while maintaining model quality.
Original source
This story was published by arXiv cs.AI and written by Qingao Yi, Jiaang Duan, Jun Zhang, Haiyan Zhao, Shiyou Qian, Dingyu Yang, Jian Cao, Jinghua Tang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


