ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
Published in Annual Meeting of the Association for Computational Linguistics (ACL), Oral, 2026
Authors: David H. Yang, Yuxuan Zhu, Mohammad Mohammadi Amiri, Keerthiram Murugesan, Tejaswini Pedapati, Subhajit Chaudhury, Pin-Yu Chen.
ZoomR targets the growing KV cache memory requirements of language models that generate long reasoning traces. It uses retrieval at multiple granularities to reduce GPU memory consumption during reasoning while preserving access to relevant cached information.
| Paper | arXiv |
