SyncAI.news, a Varaisys broadcasting
Introducing Training Cluster as a Service - a new collaboration with NVIDIA
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

Introducing Training Cluster as a Service - a new collaboration with NVIDIA

Today at GTC Paris, we are excited to announce Training Cluster as a Service in collaboration with NVIDIA, to make large GPU clusters more easily accessible for research organizations all over the world, so they can train the foundational models of tomorrow in every domain.

Making GPU Clusters Accessible

Many Gigawatt-size GPU supercluster projects are being built to train the next gen of AI models. This can make it seem that the compute gap between the “GPU poor” and the “GPU rich” is quickly widening. But the GPUs are out there, as hyperscalers, regional and AI-native cloud providers all quickly expand their capacity.

How do we then connect AI compute capacity with the researchers who need it? How do we enable universities, national research labs and companies all over the world to build their own models?

This is what Hugging Face and NVIDIA are tackling with Training Cluster as a Service - providing GPU cluster accessibility, with the flexibility to only pay for the duration of training runs.

To get started, any of the 250,000 organizations on Hugging Face can request the GPU cluster size they need, when they need it.

How it works

To get started, you can request a GPU cluster on behalf of your organization at hf.co/training-cluster

Training Cluster as a Service integrates key components from NVIDIA and Hugging Face into a complete solution:

  • NVIDIA Cloud Partners provide capacity for the latest NVIDIA accelerated computing like NVIDIA Hopper and NVIDIA GB200 in regional datacenters, all centralized within NVIDIA DGX Cloud
  • NVIDIA DGX Cloud Lepton - announced today at GTC Paris - provides easy access to the infrastructure provisioned for researchers, and enables training run scheduling and monitoring
  • Hugging Face developer resources and open source libraries make it easy to get training runs started.

Clusters at Work

Advancing Rare Genetic Disease Research with TIGEM

-- Diego di Bernardo, Coordinator of the Genomic Medicine program at TIGEM

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News