SyncAI.news, a Varaisys broadcasting
MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
IT

Itzel Tlelo-Coyotecatl, Hugo Jair Escalante

· 1 min read

ResearcharXiv cs.CL

MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos

arXiv:2609.31553v1 Announce Type: new Abstract: Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.

Original source

This story was published by arXiv cs.CL and written by Itzel Tlelo-Coyotecatl, Hugo Jair Escalante. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News