SyncAI.news, a Varaisys broadcasting
Anchored or Drifting: What Recursive Self-Generation Reveals About Training Data
W{

Wojciech {\L}apacz, Stanis{\l}aw Pawlak

· 1 min read

ResearcharXiv cs.AI

Anchored or Drifting: What Recursive Self-Generation Reveals About Training Data

arXiv:2606.31991v2 Announce Type: replace-cross Abstract: Large generative models are known to memorize their training data, posing severe privacy risks. Yet, current methods to detect training membership typically rely on the weak signals of a single forward pass. In this work, we find that training samples and unseen (held-out) data follow visibly different trajectories under recursive self-generation -- repeatedly feeding a model's output back as its next input. Held-out samples \emph{drift}: they lose the specifics of the original within a few steps. Training samples stay \emph{anchored}, degrading far more slowly. We show that the membership signal this produces holds across model modalities, architectures and scales, spanning language, diffusion, and autoregressive vision models. Furthermore, these recursive trajectories provide a signal that raises membership inference TPR at $1\%$ FPR for nearly every attack we evaluate, roughly doubling it on the weakest baselines and still improving the strongest, which shows that a model's behavior under recursion carries membership evidence that a single query does not.

Original source

This story was published by arXiv cs.AI and written by Wojciech {\L}apacz, Stanis{\l}aw Pawlak. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News