
TD
Thao Do, Dinh Phu Tran, An Vo, Seon Kwon Kim, Daeyoung Kim
· 1 min read
ResearcharXiv cs.CL
EnComp: Lightweight Encoder-Only Context Compression for Retrieval-Augmented Question Answering
arXiv:2603.09222v2 Announce Type: replace
Abstract: Efficient context compression is critical for retrieval-augmented question answering in resource-constrained settings, where long retrieved contexts increase latency, memory use, and LLM reader cost. We propose a lightweight encoder-only framework for query-driven sentence pruning that preserves answer-critical evidence while aggressively reducing irrelevant context. Our method learns marginal contribution scores for sentences using counterfactual training signals and optimizes a contrastive ranking objective that separates critical evidence from noncritical context. Our approach scores all sentences from a single full-context encoding, enabling fast inference with low computational overhead. Experiments show that it maintains accuracy comparable to the strongest baseline while using 3.7$\times$ less peak memory and achieving nearly 3$\times$ lower compression latency, demonstrating an effective quality--efficiency trade-off for practical resource-constrained deployment.
Original source
This story was published by arXiv cs.CL and written by Thao Do, Dinh Phu Tran, An Vo, Seon Kwon Kim, Daeyoung Kim. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


