
JW
Jiayun Wu, Peng Zhang, Yuanyuan Lu, Shan Qu, Ning Gu, Tun Lu
· 1 min read
ResearcharXiv cs.CL
Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
arXiv:2603.20212v2 Announce Type: replace
Abstract: Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior accuracy through chain-of-thought (CoT) reasoning, they incur substantial computational costs. Conversely, Scalar Reward Models (SRMs) offer efficiency but suffer from limited performance and adaptability in complex scenarios. We introduce Fast-Slow Thinking Reward Models (F/S-RM), a hybrid RM architecture inspired by Dual Process Theory. It trains a single model to integrate two distinct reward paradigms: scalar-style first-token pairwise judgment (fast thinking) and CoT-based judgment (slow thinking), regulated by a dual-confidence activation mechanism that determines when to activate slow thinking. Under hybrid inference, F/S-RM achieves state-of-the-art accuracy with an average score of 84.3 across benchmarks, while reducing token consumption by 22.5%.
Original source
This story was published by arXiv cs.CL and written by Jiayun Wu, Peng Zhang, Yuanyuan Lu, Shan Qu, Ning Gu, Tun Lu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


