Fine-grained KV cache reuse that assembles query-relevant nugget caches from retrieved chunks through two-stage retrieval, improving the accuracy-latency Pareto frontier for long-context RAG
Bridging the training-inference gap for dense phrase retrieval via unified loss and hard negatives mined through efficient subcorpus validation - ___[Findings of EMNLP 2022](https://2022.emnlp.org)___ · ___[SustaiNLP @ EMNLP 2022](https://sites.google.com/view/sustainlp2022)___ · ___[KRLM @ ICML 2022](https://knowledge-retrieval-workshop.github.io)___
Dynamic sequence length reduction to further enhance the inference efficiency of TinyBERT beyond static compression - ___[ENLSP Workshop @ NeurIPS 2021](https://neurips2021-nlp.github.io)___
Train-once, anytime-inference framework for any transformer via length drop training and multi-objective evolutionary search - ___[ACL 2021](https://2021.aclweb.org)___ · ___[SustaiNLP @ EMNLP 2021](https://sites.google.com/view/sustainlp2021)___ (Best Paper Award)