
TC
Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen
· 1 min read
ResearcharXiv cs.CV
Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation
arXiv:2609.19729v1 Announce Type: new
Abstract: Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response $R( \Delta, \, t_{\text{denoise}})$, a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from $R$, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.
Original source
This story was published by arXiv cs.CV and written by Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


