SyncAI.news, a Varaisys broadcasting
Beyond Math and Code: Lightweight Corpus-Grounded Process Rewards for Factual Question Answering
SF

Shicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Hanrong Zhang, Lingwei Wei, Henry Peng Zou, Chengquan Guo, Jie Yang, Honghui Bao, Zhiwei Liu, Lu Cheng, Philip S. Yu

· 1 min read

ResearcharXiv cs.CL

Beyond Math and Code: Lightweight Corpus-Grounded Process Rewards for Factual Question Answering

arXiv:2605.29648v2 Announce Type: replace Abstract: Process supervision during reinforcement learning (RL) post-training matters in factual question answering (QA) because responses receiving positive response-level rewards can still contain sentence-level factual errors. Unlike math and code, factual QA lacks inexpensive programmatic checks, making reward computation a bottleneck. Existing methods rely on neural verifiers to score individual sentences, requiring extensive model inference as RL repeats these checks across many sampled responses at every update. We therefore introduce CorVer (Corpus Verify), a lightweight training-time approach to process supervision that derives sentence-level rewards from corpus co-occurrence statistics. A 0.5B extractor identifies subject-object pairs, indexed corpus queries supply their co-occurrence counts, and the resulting sentence rewards are assigned to the corresponding tokens for RL. On Qwen3-4B and Qwen3-8B, CorVer reduces mean complete training time by 5.5-10.4 times relative to the four factuality-RL baselines. Across the four models evaluated against these baselines, CorVer achieves the highest accuracy in 17 of 20 model-benchmark settings. CorVer outperforms the unmodified models in all 30 standard factual-QA settings (six models from three families across five benchmarks) and remains effective on two additional multi-hop QA datasets.

Original source

This story was published by arXiv cs.CL and written by Shicheng Fan, Haochang Hao, Dehai Min, Weihao Liu, Hanrong Zhang, Lingwei Wei, Henry Peng Zou, Chengquan Guo, Jie Yang, Honghui Bao, Zhiwei Liu, Lu Cheng, Philip S. Yu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News