EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Abstract
Next-generation visual assistants such as smart glasses, embodied agents, and always-on life-logging systems must reason over an entire day or more of continuous visual experience. In such ultra-long video settings, relevant information is sparsely distributed across hours or days, making this mainly a memory problem: models must accumulate information over time, recall previously observed states, track temporal order, and abstract recurring patterns from past experience. However, existing week-long video benchmarks are still primarily designed for perception and recognition, such as locating a specific moment or summarizing global content, rather than reasoning that requires accumulating and integrating evidence across multiple days. To address this gap, we introduce EgoMemory, a comprehensive benchmark that systematically evaluates week-long egocentric video understanding through the lens of memory. EgoMemory evaluates three complementary memory types: entity memory, tracking how object states evolve and change across days; event memory, recalling and ordering activities separated by hours or days; and behavior memory, abstracting recurring patterns from sparse, repeated observations over the whole week period. EgoMemory comprises 500 questions across three memory types and six core challenges, with an average of 4.1 pieces of visual evidence per question requiring memory backtracking over 26 hours. We evaluate EgoMemory on 13 methods across MLLMs and agentic frameworks, revealing that even the best model achieves only 36.8% overall accuracy. Further analysis shows that performance degrades as evidence spans longer temporal horizons, revealing that long-horizon memory remains far from solved. We also find that neither much-denser frame sampling nor auxiliary captions yield consistent gains, suggesting the core bottleneck lies in how models store and retrieve information over long temporal horizons. We hope this challenging benchmark establishes a strong foundation for evaluating and advancing long-context multimodal memory systems.