
QV
Quyet V. Do, Thinh Pham, Nguyen Nguyen, Sha Li, Pratibha Zunjare, Tu Vu
· 1 min read
ResearcharXiv cs.CL
$\pi^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
arXiv:2604.05114v2 Announce Type: replace
Abstract: We study a QA curation pipeline for improving long-context complex reasoning in large language models (LLMs). Our approach, $\pi^2$, constructs high-quality reasoning data through rigorous QA curation: 1) extracting and expanding tables from Wikipedia, 2) from the collected tables together with relevant metadata, generating complex reasoning questions whose answers are automatically determined and validated through dual-path code execution, 3) finally, back-translating chain-of-thoughts solutions grounded in realistic context. Supervised fine-tuning with gpt-oss-20b and Qwen3-4B-Instruct-2507 on $\pi^2$ yields consistent improvements across four long-context reasoning benchmarks and our alike $\pi^2$-Bench, with average absolute accuracy gains of +6.25% and +3.37% respectively. Through deeper analyses, we observe that reasoning style contributes little, while faithful reasoning patterns discovered by back translation and grounded realistic long context, as $\pi^2$ is designed for, are crucial for the improvement. Our code, data, and models are fully open-source at https://github.com/vtpss/pi-squared.
Original source
This story was published by arXiv cs.CL and written by Quyet V. Do, Thinh Pham, Nguyen Nguyen, Sha Li, Pratibha Zunjare, Tu Vu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


