OjaKV: Context-Aware Online Low-Rank KV Cache Compression

Published in Findings of the Association for Computational Linguistics (ACL), 2026

Authors: Yuxuan Zhu, David H. Yang, Mohammad Mohammadi Amiri, Keerthiram Murugesan, Tejaswini Pedapati, Pin-Yu Chen.

OjaKV reduces the memory cost of long-context language model inference through adaptive low-rank KV cache compression. It retains important initial and recent tokens at full rank while updating the compression subspace online with Oja’s algorithm to follow changes in context. The method works with FlashAttention and requires no model fine-tuning.

PaperarXiv