
TA
Thea Aviss
· 1 min read
ResearcharXiv cs.CL
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
arXiv:2605.00206v2 Announce Type: replace-cross
Abstract: Current transformers discard their rich latent residual stream between positions, reconstructing latent reasoning context at each new position and leaving potential reasoning capacity untapped. The State Stream Transformer (SST) V2 enables parameter-efficient reasoning in continuous latent space through an FFN-driven nonlinear recurrence at each decoder layer, where latent states are streamed horizontally across the full sequence via a learned blend. This same mechanism supports continuous latent deliberation per position at inference time, dedicating additional FLOPs to exploring abstract reasoning before committing to a token. A two-pass parallel training procedure approximates the sequential recurrence, making co-training computationally practical. Hidden state analysis shows that the state stream facilitates reasoning through sharp, content-dependent reorganisations in continuous latent space; the LM head exposes the resulting latent belief states through the output distribution, while the state stream carries them forward to influence future positions. A learned probe shows that at the first generated token, the latent state already predicts whether the eventual answer will survive or break under additional latent computation for every subsequent position. Co-trained into an existing 27B backbone using only a small dataset of GSM8K examples and evaluated using an oracle to route each question to a depth of one to four recurrent forward passes per token, the SST achieves an architectural capacity bound of 61.11% on out-of-distribution GPQA-Diamond, a +15.15 point gain over a fine-tuning-matched baseline, and cuts that same baseline's remaining GSM8K errors by 46%. Together, these results provide a method for efficiently training a nonlinear recurrence and show that the state stream provides additional reasoning capacity beyond that of an otherwise training-matched transformer.
Original source
This story was published by arXiv cs.CL and written by Thea Aviss. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


