
SG
Shuyu Guo, Shuo Zhang, Zhaochun Ren
· 1 min read
ResearcharXiv cs.CL
Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG
arXiv:2609.05152v2 Announce Type: replace
Abstract: Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this cost by encoding documents into compact continuous representations that can be precomputed and reused across queries. However, many existing methods train compressed models by distilling from a full-context teacher. When the teacher is wrong, such distillation can reinforce its errors, while teacher imitation provides no direct signal for improving beyond the teacher. We propose DEX-Comp, a two-stage training recipe that separates reliable imitation from targeted exploration. Pure Distillation learns only from teacher-correct questions to mitigate error propagation, while Hard Exploration applies outcome-based reinforcement learning to teacher-failed questions to directly optimize answer correctness. Across five open-domain QA benchmarks and retrieval depths from top-$5$ to top-$30$, DEX-Comp at $16\times$ compression outperforms all evaluated compression baselines and surpasses the untuned full-context RAG model in average accuracy, while reducing time-to-first-token by $4.4\times$--$23.7\times$. Evaluations across additional datasets and backbones further demonstrate its generalization.
Original source
This story was published by arXiv cs.CL and written by Shuyu Guo, Shuo Zhang, Zhaochun Ren. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


