SyncAI.news, a Varaisys broadcasting
Can We Anticipate Violence? Multimodal Learning from Pre-Incident Behavioral Cues
SP

Sindhuja Penchala, Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi

· 1 min read

ResearcharXiv cs.CV

Can We Anticipate Violence? Multimodal Learning from Pre-Incident Behavioral Cues

arXiv:2609.40014v1 Announce Type: new Abstract: Detecting violence after it begins is important from recognizing behavioral cues that appear immediately beforehand. This work studies short-horizon pre-incident risk recognition from multimodal video signals. We construct a binary Normal-versus-Risky setting from temporally annotated XD-Violence clips, using 443 samples with source-level separation across training, validation, and test sets. Each sample consists of a variable-length pre-incident clip, with its duration determined by the observable behavioral context preceding the incident. The inci- dent itself is excluded from all input clips. We evaluate three complementary information sources: facial-region appearance, temporally aligned audio, and body-motion features derived from tracked keypoints. Controlled ablations are performed with Swin-Tiny, ViT-Tiny, and DeiT-Tiny to measure the contribution of each modality under the same split. Results show that combining all modalities is more effective than using any other combination alone. The best configuration, Deit-Tiny with audio, facial appearance, and motion, achieves 91.21% accuracy, 88.96% balanced accuracy, 93.65% F1-score, and 96.38% ROC-AUC on the held-out test set. These results suggest that complementary appearance, acoustic, and kinematic cues provide useful evidence for recognizing elevated pre-incident risk.

Original source

This story was published by arXiv cs.CV and written by Sindhuja Penchala, Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News