Fine-grained KV cache reuse that assembles query-relevant nugget caches from retrieved chunks through two-stage retrieval, improving the accuracy-latency Pareto frontier for long-context RAG
Hierarchical RAG framework that compresses and partitions documents at multiple scales to address imprecise retrieval and fragmented context in long-context LLMs