
SL
Shenghao Li, Lin Zhao
· 1 min read
ResearcharXiv cs.CL
DualGuard: Dual-Mode Quality Control for Logic-Preserving Data Augmentation
arXiv:2609.32431v1 Announce Type: new
Abstract: Large language models provide a practical way to generate augmented data for logical reasoning at scale, but a larger generation volume does not guarantee semantic, label, or logical reliability. Existing work has improved generation quality through generation constraints, candidate validation, filtering, and feedback-based revision; however, once a quality judgment is available, deciding whether a candidate should be retained, filtered, or repaired remains an important control problem. We propose DualGuard, a dual-mode quality-control framework for logic-preserving data augmentation. The first mode uses the current instance and candidate batch for selective retention, filtering, attribution, and targeted feedback. The second mode accumulates cross-instance execution records of augmentation actions on top of per-sample diagnosis and attribution, compares new executions against each action's own historical behavior, and supports retrospective anomaly inspection, targeted rollback, and bounded repair. Both modes share semantic verification and additionally use symbolic verification when a reliable logical form is available. Across seven downstream tasks in the Two-Stage Transfer setting, DualGuard achieves the highest Accuracy on five tasks and outperforms the no-augmentation BERT baseline on all seven. Controlled ablations further show complementary roles for Memory, Z3, and history-aware anomaly control.
Original source
This story was published by arXiv cs.CL and written by Shenghao Li, Lin Zhao. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


