SyncAI.news, a Varaisys broadcasting
Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation
AS

Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad

· 1 min read

ResearcharXiv cs.AI

Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation

arXiv:2512.13142v5 Announce Type: replace Abstract: People use LLMs for reproductive-health questions, including abortion-related support. A response can sound supportive while answers reinforce harmful assumptions: judgment is likely, secrecy is safer, and support is limited. We introduce behavioral coherence evaluation, a design-time method that uses validation evidence from an established instrument to test relations among outputs. Using the Individual Level Abortion Stigma Scale, we prompted five LLMs to complete questionnaires for 627 personas and reviewed flagged patterns with five reproductive-health experts. Models scored personas lower on self-judgment but higher on worries about judgment; most made worries the highest-scoring dimension, although it was lowest in the ILAS reference sample. Four of five models reversed the reference direction by generating significantly higher worries about judgment scores for Black personas. Models defaulted to extreme secrecy after abortion despite varying stigma patterns across personas. Expert review showed that disclosure guidance requires context about relationship safety, legal risk, and trusted support.

Original source

This story was published by arXiv cs.AI and written by Anika Sharma, Malavika Mampally, Chidaksh Ravuru, Kandyce Brennan, Neil Gaikwad. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News