SyncAI.news, a Varaisys broadcasting
Pre-train to Gain: Robust Learning Without Clean Labels
DS

David Szczecina, Nicholas Pellegrino, Paul Fieguth

· 1 min read

ResearcharXiv cs.AI

Pre-train to Gain: Robust Learning Without Clean Labels

arXiv:2511.20844v2 Announce Type: replace-cross Abstract: Training deep networks with noisy labels leads to poor generalization and degraded accuracy due to overfitting to label noise. Existing approaches for learning with noisy labels often rely on the availability of a clean subset of data. By pre-training a feature extractor on the target dataset without labels using in-domain self-supervised learning (SSL), followed by standard supervised training on the same noisy dataset, we can train a more noise robust model without requiring a subset with clean labels. We evaluate both contrastive and non-contrastive SSL pre-training methods across datasets with synthetic and real-world label noise, demonstrating the broad applicability of our approach across large-scale datasets, diverse downstream tasks, and model architectures. Across all noise rates, in-domain self-supervised pre-training consistently improves classification accuracy and downstream label-error detection (F1 and Balanced Accuracy) compared with supervised training from scratch. The performance gap widens as the noise rate increases, demonstrating improved robustness. Notably, our approach achieves comparable results to ImageNet and DinoV2 pre-trained models at low noise levels, while substantially outperforming them under high noise conditions.

Original source

This story was published by arXiv cs.AI and written by David Szczecina, Nicholas Pellegrino, Paul Fieguth. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News