SyncAI.news, a Varaisys broadcasting
Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG
SG

Shuyu Guo, Shuo Zhang, Zhaochun Ren

· 1 min read

ResearcharXiv cs.CL

Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG

arXiv:2609.05152v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves knowledge-intensive generation by conditioning language models on retrieved documents, but processing these documents becomes increasingly expensive as retrieval depth grows. Soft context compression reduces this cost by encoding documents into compact continuous representations that can be precomputed and reused across queries. However, many existing methods train compressed models by distilling from a full-context teacher. When the teacher is wrong, such distillation can reinforce its errors, while teacher imitation provides no direct signal for improving beyond the teacher. We propose DEX-Comp, a two-stage training recipe that separates reliable imitation from targeted exploration. Pure Distillation learns only from teacher-correct questions to mitigate error propagation, while Hard Exploration applies outcome-based reinforcement learning to teacher-failed questions to directly optimize answer correctness. Across five open-domain QA benchmarks and retrieval depths from top-$5$ to top-$30$, DEX-Comp at $16\times$ compression outperforms all evaluated compression baselines and surpasses the untuned full-context RAG model in average accuracy, while reducing time-to-first-token by $4.4\times$--$23.7\times$. Evaluations across additional datasets and backbones further demonstrate its generalization.

Original source

This story was published by arXiv cs.CL and written by Shuyu Guo, Shuo Zhang, Zhaochun Ren. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News