SyncAI.news, a Varaisys broadcasting
VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text
TH

Trieu Hai Nguyen, Sivaswamy Akilesh

· 1 min read

ResearcharXiv cs.CL

VietBinoculars: A Zero-Shot Approach for Detecting Vietnamese LLM-Generated Text

arXiv:2509.26189v2 Announce Type: replace Abstract: The rapid proliferation of Large Language Models has intensified the challenge of distinguishing LLM-generated text from human writing in non-English languages. This study introduces VietBinoculars, a zero-shot detection framework coupling PhoGPT-4B observer and performer models with calibrated global decision thresholds. By utilizing specialized Vietnamese BPE tokenization, the method eliminates byte-level fragmentation and probability dilution common in massive multilingual backbones. Evaluated across multi-domain benchmarks, VietBinoculars achieves an area under the ROC curve exceeding 0.99. Under optimal Youden's J thresholds and greedy decoding, detection accuracy reaches at least 98.78\%, while significantly outperforming baseline Binoculars, zero-shot detectors, and commercial tools on creative Capybara prompts. Even under a strict false positive rate constraint of 0.06\%, the detector maintains F1-scores between 83.15\% and 94.70\%. Detection performance consistently improves with sequence length, stabilizing at optimal accuracy for passages containing 450 to 550 tokens. Extended stress testing across 48 distinct model-decoding configurations and three post-generation rewriting strategies delineates practical operational boundaries. VietBinoculars exhibits robust resilience against single-pass paraphrasing and human-style revisions, but experiences notable performance degradation under high-entropy sampling and iterative double paraphrasing.

Original source

This story was published by arXiv cs.CL and written by Trieu Hai Nguyen, Sivaswamy Akilesh. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News