
AG
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
· 1 min read
ResearcharXiv cs.CL
Strategically Diverse Sampling for Self-Training
arXiv:2609.31571v1 Announce Type: new
Abstract: Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.
Original source
This story was published by arXiv cs.CL and written by Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


