SyncAI.news, a Varaisys broadcasting
Synthetic Data Characterization via Training Dynamics
IL

Irene Lago, Ana Ezquerro, David Vilares

· 1 min read

ResearcharXiv cs.LG

Synthetic Data Characterization via Training Dynamics

arXiv:2609.39447v1 Announce Type: cross Abstract: Interpreting properties of LLM-generated data is important for understanding its utility and limitations across learning tasks. In this work, we characterize synthetic data through sample-level learnability, studying variation among LLM families and scales, alongside human-written data as a reference. We first generate synthetic datasets spanning single- and multi-label classification, labeling, and tree prediction tasks. We then derive empirical data distributions from encoder training dynamics for both machine and organic data, and estimate the robustness of these distributions across encoders. Finally, we evaluate how data selection strategies based on these learnability signals affect both data sources differently.

Original source

This story was published by arXiv cs.LG and written by Irene Lago, Ana Ezquerro, David Vilares. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News