
SD
Sagar Deb, Devam Shah, Ashwanth Krishnan
· 1 min read
ResearcharXiv cs.LG
Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions
arXiv:2609.31675v1 Announce Type: new
Abstract: We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world generator with a fixed observe-intervene-submit API and a three-layer scoring contract that separates skeleton recovery, DAG recovery, and intervention efficiency. On the current six-level ladder, PC with a greedy active orientation heuristic is the strongest non-oracle method (directed F1 42.7%, SHD 4.79), ahead of Claude Sonnet 4.6 raw active (31.7%, 7.25) and GPT-5.4 raw active (22.9%, 9.27). The most informative diagnostic is the precision-recall decomposition: PC under-commits with high precision, LLMs over-commit with lower precision, and statistical-tool access often increases abstention rather than useful intervention. A structure-blind random DAG baseline reaches 23.6% directed F1 on this dense v0 ladder; a density probe lowers this floor to 16.9%, motivating the v1 calibration pass. The current results should therefore be read as a benchmark audit and calibration report, not as evidence that current LLMs solve active causal discovery.
Original source
This story was published by arXiv cs.LG and written by Sagar Deb, Devam Shah, Ashwanth Krishnan. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


