InfMem: Learning System-2 Memory Control for Long-Context Agent
Xinyu Wang ⋅ Peng Lu ⋅ Xiao-Wen Chang ⋅ Lifeng Shang ⋅ Jinpeng Li ⋅ Fei Mi ⋅ Prasanna Parthasarathi ⋅ Yufei Cui ⋅ 明泽 李
Abstract
Reasoning over ultra-long documents requires synthesizing sparse evidence scattered across distant segments under strict memory constraints. While streaming agents enable scalable processing, their passive memory update strategy often fails to preserve low-salience \emph{bridging evidence} required for multi-hop reasoning. We propose \textbf{InfMem}, a control-centric agent that instantiates System-2-style control via a \textsc{PreThink--Retrieve--Write} protocol. InfMem actively monitors evidence sufficiency, performs targeted in-document retrieval, and applies evidence-aware joint compression to update a bounded memory. To ensure reliable control, we introduce a practical SFT$\rightarrow$RL training recipe that aligns retrieval, writing, and stopping decisions with end-task correctness. On ultra-long QA benchmarks from 32k to 1M tokens, InfMem consistently outperforms MemAgent across backbones. Specifically, InfMem improves average absolute accuracy by \textbf{+10.17}, \textbf{+11.84}, and \textbf{+8.23} points on Qwen3-1.7B, Qwen3-4B, and Qwen2.5-7B, respectively, while reducing inference time by \textbf{3.9$\times$} on average (up to 5.1$\times$) via adaptive early stopping.
Successful Page Load