
IZ
Itai Zehavi, Fanny Jourdan, Ulrich Aivodji
· 1 min read
ResearcharXiv cs.AI
Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders
arXiv:2609.31056v1 Announce Type: new
Abstract: Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive training promotes target-selective features and that our selection score controls expected background knowledge perturbation. We validate SCALPEL experimentally on TOFU across Qwen, Llama, and Gemma, where it substantially improves over NMF and standard SAE interventions and is competitive with Gradient Difference and RMU, bridging mechanistic interpretability and fine-grained unlearning.
Original source
This story was published by arXiv cs.AI and written by Itai Zehavi, Fanny Jourdan, Ulrich Aivodji. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


