
ZL
Zewei Li, Qiaoqiao Ren, Hang Yang, S. C. Wong, Stergios-Aristoteles Mitoulis, Yun Ye
· 1 min read
ResearcharXiv cs.LG
Not All Synthetic Data Are Equal: Expert-Committee Audit Screening for Imbalanced Crash-Injury-Severity Prediction in Automated Driving Systems
arXiv:2609.29687v1 Announce Type: new
Abstract: Automated driving systems (ADSs) are increasingly operating on public roads, raising safety concerns, yet reliable prediction of crash injury severity remains difficult because crash reports are limited, severe outcomes are rare, and injury classes are highly imbalanced. Existing augmentation methods mainly increase minority-class sample size but rarely assess whether generated samples are credible for safety-critical prediction. This study proposes Expert-Committee Audit Screening (ECAS), a credibility-aware sample acceptance framework for ADS crash injury severity prediction under data imbalance. Using 1,477 incident-level ADS crashes from the National Highway Traffic Safety Administration Standing General Order records, ECAS audits generated minority samples through a real-data-only expert committee based on label support, boundary separation, committee agreement, and local plausibility. Within-class percentile normalization and Pareto non-dominated sorting select accepted samples without manually assigned evidence weights. With a fixed backbone combining normalizing flow augmentation and a Tabular Prior-data Fitted Network (TabPFN) classifier, the best ECAS configuration achieved the highest balanced accuracy, macro-F1, and minor-injury recall among all evidence configurations. Local neighborhood analysis showed that ECAS-accepted samples were better supported by nearby real minority crashes than unscreened retained samples. Shapley additive explanations and partial dependence plots further indicated that lower injury severity classes were mainly associated with crash counterpart and pre-crash movement, whereas moderate-plus injuries were more sensitive to posted speed limit and operating context. These findings support a shift from quantity-oriented augmentation to credibility-aware sample acceptance for ADS safety prediction and risk governance.
Original source
This story was published by arXiv cs.LG and written by Zewei Li, Qiaoqiao Ren, Hang Yang, S. C. Wong, Stergios-Aristoteles Mitoulis, Yun Ye. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


