
HC
Hongren Chen, Jiayang He
· 1 min read
ResearcharXiv cs.LG
Low-Bit Recurrent States in Hybrid Language Models
arXiv:2609.30950v1 Announce Type: new
Abstract: Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative log-likelihood by factors of 3.3--27.9 relative to the best of seven baselines across three hybrid models; metadata costs vary. At six bits, negative log-likelihood differs from the FP32-state baseline by less than 0.005 nats. Ablations separate gains from variable bit widths, decay weighting, and range normalization. With less frequent write-backs, gains diminish and depend on the model and budget.
Original source
This story was published by arXiv cs.LG and written by Hongren Chen, Jiayang He. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


