
MU
Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov
· 1 min read
ResearcharXiv cs.CL
Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark
arXiv:2609.28090v1 Announce Type: new
Abstract: Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0\% closed and clean-aware code recall, but open prompts over-flag 93.8\% of clean code controls, and clean-aware all-three specificity is 87.5\% even where recall saturates. A clean-aware warning drops DeepSeek code false positives from 20.8\% (95\% CI 11.7--34.3) to 0.0\% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.
Original source
This story was published by arXiv cs.CL and written by Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


