SyncAI.news, a Varaisys broadcasting
Scaling-up BERT Inference on CPU (Part 1)
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

Scaling-up BERT Inference on CPU (Part 1)

1. Context and Motivations

Back in October 2019, my colleague Lysandre Debut published a comprehensive (at the time) inference performance benchmarking blog (1).

Since then, 🤗 transformers (2) welcomed a tremendous number of new architectures and thousands of new models were added to the 🤗 hub (3) which now counts more than 9,000 of them as of first quarter of 2021.

As the NLP landscape keeps trending towards more and more BERT-like models being used in production, it remains challenging to efficiently deploy and run these architectures at scale.
This is why we recently introduced our 🤗 Inference API: to let you focus on building value for your users and customers, rather than digging into all the highly technical aspects of running such models.

This blog post is the first part of a series which will cover most of the hardware and software optimizations to better leverage CPUs for BERT model inference.

For this initial blog post, we will cover the hardware part:

  • Setting up a baseline - Out of the box results
  • Practical & technical considerations when leveraging modern CPUs for CPU-bound tasks
  • Core count scaling - Does increasing the number of cores actually give better performance?
  • Batch size scaling - Increasing throughput with multiple parallel & independent model instances

We decided to focus on the most famous Transformer model architecture, BERT (Delvin & al. 2018) (4). While we focus this blog post on BERT-like models to keep the article concise, all the described techniques can be applied to any architecture on the Hugging Face model hub.
In this blog post we will not describe in detail the Transformer architecture - to learn about that I can't recommend enough the Illustrated Transformer blogpost from Jay Alammar (5).

Today's goals are to give you an idea of where we are from an Open Source perspective using BERT-like models for inference on PyTorch and TensorFlow, and also what you can easily leverage to speedup inference.

2. Benchmarking methodology








Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News