
JC
Jacob Carlson
· 1 min read
ResearcharXiv cs.LG
Interpretable Discovery from Unstructured Data: A High-Dimensional Approach
arXiv:2511.01680v5 Announce Type: replace-cross
Abstract: We propose an automatic, general-purpose framework for making discoveries from unstructured data (e.g., text data from open-ended surveys of economic beliefs). The framework leverages recent methods from the literature on AI interpretability to transform unstructured datasets into high-dimensional, structured datasets of interpretable concept measurements; specifies concept-level parameters and null hypotheses based on this transformed dataset; tests these hypotheses using algorithms validated by new results in high-dimensional multiple testing, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is robust to data snooping, mitigates under-exploration, and facilitates fast and inexpensive sensitivity analysis and replication. We revisit applications to recent descriptive and causal analyses of unstructured data in empirical economics, and find this framework is able to automatically replicate existing discoveries, add nuance to others, and make entirely new discoveries as well.
Original source
This story was published by arXiv cs.LG and written by Jacob Carlson. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


