ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval

Published in Annual Meeting of the Association for Computational Linguistics (ACL), Oral, 2026

Authors: David H. Yang, Yuxuan Zhu, Mohammad Mohammadi Amiri, Keerthiram Murugesan, Tejaswini Pedapati, Subhajit Chaudhury, Pin-Yu Chen.

ZoomR targets the growing KV cache memory requirements of language models that generate long reasoning traces. It uses retrieval at multiple granularities to reduce GPU memory consumption during reasoning while preserving access to relevant cached information.

PaperarXiv