SyncAI.news, a Varaisys broadcasting
CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
M\

Marko \v{R}eh\'a\v{c}ek, V\'it\v{e}zslav Du\v{s}ek, Martin Rusinko, V\'it Nov\'a\v{c}ek

· 1 min read

ResearcharXiv cs.CL

CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings

arXiv:2609.31062v1 Announce Type: new Abstract: Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential linear direction. To train the probes, we cluster 79,658 Czech oncology search queries with BERTopic and use these clusters to generate pairs of queries with contrastive risk levels via few-shot prompting. We evaluate the approach on 200 queries (90 real, 110 synthetic), each graded by two oncologists, against two open-weight LLMs and a frontier LLM. We find that urgency-based axes are recoverable as linear directions, and the probes are competitive with open-weight LLMs (no significant differences in quadratic-weighted kappa) at a fraction of the latency. Each axis yields a scalar score that clinicians can inspect and use to set escalation thresholds. The pipeline requires only search logs, axis definitions, and black-box access to the embedding model, suggesting transferability across healthcare domains. Robust validation on new queries and axes remains future work.

Original source

This story was published by arXiv cs.CL and written by Marko \v{R}eh\'a\v{c}ek, V\'it\v{e}zslav Du\v{s}ek, Martin Rusinko, V\'it Nov\'a\v{c}ek. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News