
YZ
Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
· 1 min read
ResearcharXiv cs.CL
CausalVerify: End-to-End Verification of Causal Analyses by Language Models
arXiv:2609.07944v2 Announce Type: replace-cross
Abstract: Language models increasingly perform empirical analyses end to end, yet existing evaluations assess the written explanation or whether generated code executes, not whether the executed workflow recovers the intended causal estimand. We introduce CausalVerify, an execution-grounded benchmark for end-to-end causal analysis that follows a model from research-context interpretation to estimand recovery. It scores this workflow at four distinct layers: method recognition, design specification, executable implementation, and estimand recovery. It combines 259 real-paper contexts, 100 fixed-seed synthetic scenarios with executable reference estimates, and 23 paper-twin pairs in which a model commits to a design before seeing the data and its executed analysis is scored against a canonical estimator on the same realised dataset. Execution is not correctness. Among 426 model-written workflows that run without error, 15.5% fail verification, and a keyword score of the effect direction stated in the text is a poor proxy for recovery. In a 23-pair, nine-model paired study, replacing a model's committed design with the reference design and its execution conventions raises joint recovery of the point estimate and standard error from 15.0% to 51.5%; yet 48.5% of eligible seeds still fail under the reference design. The direction replicates on six pairs built afterwards under a frozen construction protocol, although on the two newest pairs the gain is confined to models from the family that built the references. Design specification is consequential but not sufficient: a plausible method and runnable code do not guarantee recovery, and even supplying the reference design and its conventions leaves substantial downstream failure. CausalVerify evaluates the executed workflow rather than its surface plausibility, and every reported number is recomputed from frozen artifacts by a single script.
Original source
This story was published by arXiv cs.CL and written by Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


