SyncAI.news, a Varaisys broadcasting
ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
MM

Mehdi Makni, Ryan Lucas, Rahul Mazumder

· 1 min read

ResearcharXiv cs.AI

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

arXiv:2609.36120v1 Announce Type: cross Abstract: Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as DartQuant, but both remain hard to scale to the largest architectures. To address the computational bottlenecks in gradient-free rotation learning, we introduce two ideas for efficiency, (i) a data selection procedure which reduces the required number of calibration data points, and (ii) an exact reduction of the associated optimization on this reduced calibration set. Our data selection procedure exploits the geometric structure of the convex hull of the activations. Using this idea, we show that a carefully selected calibration set with several orders of magnitude fewer activations than state-of-the-art rotation-based methods can match their performance in low-bit quantization settings. Under this extreme data efficiency, the selected activations span an $r$-dimensional subspace with $r

Original source

This story was published by arXiv cs.AI and written by Mehdi Makni, Ryan Lucas, Rahul Mazumder. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News