RAG

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Fine-grained KV cache reuse that assembles query-relevant nugget caches from retrieved chunks through two-stage retrieval, improving the accuracy-latency Pareto frontier for long-context RAG

MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG

Hierarchical RAG framework that compresses and partitions documents at multiple scales to address imprecise retrieval and fragmented context in long-context LLMs