SyncAI.news, a Varaisys broadcasting
Interpretable Discovery from Unstructured Data: A High-Dimensional Approach
JC

Jacob Carlson

· 1 min read

ResearcharXiv cs.LG

Interpretable Discovery from Unstructured Data: A High-Dimensional Approach

arXiv:2511.01680v5 Announce Type: replace-cross Abstract: We propose an automatic, general-purpose framework for making discoveries from unstructured data (e.g., text data from open-ended surveys of economic beliefs). The framework leverages recent methods from the literature on AI interpretability to transform unstructured datasets into high-dimensional, structured datasets of interpretable concept measurements; specifies concept-level parameters and null hypotheses based on this transformed dataset; tests these hypotheses using algorithms validated by new results in high-dimensional multiple testing, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is robust to data snooping, mitigates under-exploration, and facilitates fast and inexpensive sensitivity analysis and replication. We revisit applications to recent descriptive and causal analyses of unstructured data in empirical economics, and find this framework is able to automatically replicate existing discoveries, add nuance to others, and make entirely new discoveries as well.

Original source

This story was published by arXiv cs.LG and written by Jacob Carlson. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News