
KD
Kirthi Devleker
· 1 min read
EngineeringNVIDIA Technical Blog
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token...
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs. NVIDIA GB300 NVL72 set a world record for pre-training DeepSeek-V3 671B at 1,648 TFLOPs per GPU, showing how advances across the entire AI…
Source
Original source
This story was published by NVIDIA Technical Blog and written by Kirthi Devleker. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on developer.nvidia.com


