SyncAI.news, a Varaisys broadcasting
Serverless Inference with Hugging Face and NVIDIA NIM
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

Serverless Inference with Hugging Face and NVIDIA NIM

Update: This service is deprecated and no longer available as of April 10th, 2025. For an alternative, you should consider Inference Providers

Today, we are thrilled to announce the launch of Hugging Face NVIDIA NIM API (serverless), a new service on the Hugging Face Hub, available to Enterprise Hub organizations. This new service makes it easy to use open models with the accelerated compute platform, of NVIDIA DGX Cloud accelerated compute platform for inference serving. We built this solution so that Enterprise Hub users can easily access the latest NVIDIA AI technology in a serverless way to run inference on popular Generative AI models including Llama and Mistral, using standardized APIs and a few lines of code within the Hugging Face Hub.

Serverless Inference powered by NVIDIA NIM

This new experience builds on our collaboration with NVIDIA to simplify the access and use of open Generative AI models on NVIDIA accelerated computing. One of the main challenges developers and organizations face is the upfront cost of infrastructure and the complexity of optimizing inference workloads for LLM. With Hugging Face NVIDIA NIM API (serverless), we offer an easy solution to these challenges, providing instant access to state-of-the-art open Generative AI models optimized for NVIDIA infrastructure with a simple API for running inference. The pay-as-you-go pricing model ensures that you only pay for the request time you use, making it an economical choice for businesses of all sizes.

NVIDIA NIM API (serverless) complements Train on DGX Cloud, an AI training service already available on Hugging Face.

How it works

Running serverless inference with Hugging Face models has never been easier. Here's a step-by-step guide to get you started:

Note: You need access to an Organization with a Hugging Face Enterprise Hub subscription to run Inference.

Before you begin, ensure you meet the following requirements:

Create a Fine-Grained Token

Find your NIM

Send your requests

Supported Models

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News