
KZ
Kaiyuan Zhang, Chuan Wang, Joey Zhong, Paul Fryzel, Kyle Polley, Jerry Ma, Ninghui Li
· 1 min read
ResearcharXiv cs.CL
PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations
arXiv:2609.22200v1 Announce Type: new
Abstract: LLM assistants and agentic systems log long multi-turn conversations. AI providers often scan these conversations for Personally Identifiable Information (PII) and mask the PII before storing or processing conversation data. Yet most PII detectors and benchmarks target self-contained records rather than cross-turn evaluation. To evaluate PII detection across turns in multi-turn conversations, we introduce PII-TRACE (Tracing Recurring PII Across Conversational Exchanges), to our knowledge the first PII benchmark to assess whether detectors identify PII in conversational contexts and cover every mention of a recurring identifier across turns. PII-TRACE contains 13,148 synthetic multi-turn dialogues in 13 languages with character-level spans and identifier clusters. Across eleven baselines, including frontier LLMs, no detector achieves full entity-level coverage without substantial false positives on PII-free conversations, and single-pass reading loses a third of the gold characters on long dialogues. To close this gap, we introduce PII-Tracer, a compact 0.6B-parameter detector trained with conversation-level supervision. PII-Tracer attains the highest entity-level coverage of any system we evaluate and also performs strongly on standard single-record benchmarks.
Original source
This story was published by arXiv cs.CL and written by Kaiyuan Zhang, Chuan Wang, Joey Zhong, Paul Fryzel, Kyle Polley, Jerry Ma, Ninghui Li. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


