SyncAI.news, a Varaisys broadcasting
ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights
JM

Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler

· 1 min read

ResearcharXiv cs.LG

ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights

arXiv:2609.38521v1 Announce Type: new Abstract: We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker product to the empirical Fisher information matrix of a small calibration set by Kullback--Leibler minimization, forming a Mahalanobis reconstruction loss from the result. The continuous ADMM updates from NanoQuant become solutions to Sylvester equations, while its discrete projection and deployment format remain unchanged. Because the curvature is local to a given set of weights, ShamAN-Q re-measures the input curvature statistic for each layer immediately before layer factorization, periodically refreshing all statistics on the partially quantized model. ShamAN-Q also redistributes the uniform rank from NanoQuant across layers at the same total number of bits. On Qwen3-Base, ShamAN-Q lowers WikiText-2 perplexity at $\approx$1 bpw from 27.56 to 22.96 (0.6B), 19.21 to 16.72 (1.7B), and 14.29 to 13.80 (4B) while matching or improving zero-shot accuracy on the Eleuther LM Evaluation Harness. On 0.6B, ShamAN-Q at $\approx$0.8 bpw matches the published perplexity of NanoQuant at $\approx$1.0 bpw.

Original source

This story was published by arXiv cs.LG and written by Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News