SyncAI.news, a Varaisys broadcasting
Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction
KM

Kingsuk Maitra, Shagun Sood Morteza Hosseini, Suman Gunnala, Vikram Gupta

· 1 min read

ResearcharXiv cs.LG

Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction

arXiv:2609.32319v1 Announce Type: new Abstract: We circumvent the Sanford-Hsu-Telgarsky (SHT) single-layer induction obstruction within a linear, one-step, causal, bilinear, symplectically consistent design class on the post-RoPE substrate, by lifting attention onto a symplectic phase space, mirroring Hairer's lift of Stormer-Verlet. The lift exits the premise of the SHT counting argument rather than the bound itself. The reframing exhibits the obstruction as a filter-order gap: a one-layer bilinear score realises a $z$-transform of joint order $(0,0)$, whereas the induction discriminator requires key-side order $\geq 1$. Applying the symplectic upper shear $M_\gamma:(q,p)\mapsto(q+\gamma p,p)$ to the post-RoPE query and key streams closes it. We prove this lift is unique within the factorised subclass, exactly symplectic at operator level, and requires post-RoPE placement; and in an explicit $T_4$-only Gaussian reduction we derive a closed-form two-branch induction phase transition, held out at $r=0.9876$ with zero fitted parameters. That law is an analytically solvable limit, not a robust prediction: restoring the $T_3$ channel moves $\gamma_c$ at $d_k=64$ from 1.030 to 0.569 and removes the crossover. Deployability follows by exact derivation: the KV cache is unchanged, prefix reuse and speculative decoding are preserved, overhead is $6d$ FLOPs per token per layer, INT8 headroom grows by at most $\log_2(1+2\gamma)$ bits, fused kernels are unmodified, and no parameters are added. At 91.3M parameters a supercritical sweep locates an emergence band: induction forms 3/3 seeds at $\gamma=0.80$ in a mean of 717 steps, against 2/3 seeds and 2700 steps at $\gamma=0$. Adverse results are reported as directly: a key-only half-lift reaches 0.949 against 0.811 for the symmetric operator, so if induction accuracy is the objective, the half-lift is the better construction. Forty-one notebooks and result files ship as ancillary material.

Original source

This story was published by arXiv cs.LG and written by Kingsuk Maitra, Shagun Sood Morteza Hosseini, Suman Gunnala, Vikram Gupta. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News