SyncAI.news, a Varaisys broadcasting
Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval

We introduce the concept of embedding quantization and showcase their impact on retrieval speed, memory usage, disk space, and cost. We'll discuss how embeddings can be quantized in theory and in practice, after which we introduce a demo showing a real-life retrieval scenario of 41 million Wikipedia texts.

Table of Contents

  • Why Embeddings?
    • Embeddings may struggle to scale
  • Improving scalability
    • Binary Quantization
      • Binary Quantization in Sentence Transformers
      • Binary Quantization in Vector Databases
    • Scalar (int8) Quantization
      • Scalar Quantization in Sentence Transformers
      • Scalar Quantization in Vector Databases
    • Combining Binary and Scalar Quantization
    • Quantization Experiments
    • Influence of Rescoring
      • Binary Rescoring
      • Scalar (Int8) Rescoring
      • Retrieval Speed
    • Performance Summarization
    • Demo
    • Try it yourself
    • Future work:
    • Acknowledgments
    • Citation
    • References

Why Embeddings?

Embeddings are one of the most versatile tools in natural language processing, supporting a wide variety of settings and use cases. In essence, embeddings are numerical representations of more complex objects, like text, images, audio, etc. Specifically, the objects are represented as n-dimensional vectors.

After transforming the complex objects, you can determine their similarity by calculating the similarity of the respective embeddings! This is crucial for many use cases: it serves as the backbone for recommendation systems, retrieval, one-shot or few-shot learning, outlier detection, similarity search, paraphrase detection, clustering, classification, and much more.

Embeddings may struggle to scale

However, embeddings may be challenging to scale for production use cases, which leads to expensive solutions and high latencies. Currently, many state-of-the-art models produce embeddings with 1024 dimensions, each of which is encoded in float32, i.e., they require 4 bytes per dimension. To perform retrieval over 250 million vectors, you would therefore need around 1TB of memory!

Improving scalability

or

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News