A Lightweight Extractive Compression Pipeline for Long-Context Retrieval-Augmented Generation
Abstract
Long contexts in Retrieval-Augmented Generation (RAG) cause high latency and computational costs. We propose a lightweight, extractive sentence-level summarization method that extracts the top-k relevant sentences and expands them with a dynamic window (w) of adjacent sentences to preserve semantic flow. Evaluated on both single-hop and multi-hop QA datasets, our approach with a minimal window (w = 1) successfully mitigates semantic fragmentation and significantly reduces token consumption. On single-hop tasks, it achieves an optimal compression ratio of 27.5x, with extreme configurations reaching up to 73x while outperforming the full-document baseline. On multi-hop tasks, the method provides a highly efficient trade-off, maintaining competitive accuracy while drastically cutting computational overhead. While the method provides an efficient trade-off, it remains slightly behind the full-context baseline in multi-hop tasks due to the complexity of fact-chaining.
Keywords
Retrieval-augmented generation, extractive summarization, context compression, large language models.