CortexMem: Brain-Guided Event Memory for Persistent Context in Language Models
Abstract
Long-context language models can receive large histories, but they still need a policy for deciding which past information stays active. CortexMem tests that policy choice with an external probe: held-out fMRI responses from a listener hearing natural stories. At each story point, a memory policy builds a fixed-budget context from prior words. Qwen/Qwen2.5-0.5B reads that context, and a ridge encoding model tests whether the resulting hidden state predicts held-out brain responses. The current Level C pilot uses one subject, three stories, leave-one-story-out folds, HRF lags 1,2,3,4, and seven policies: ShortWindow, FixedSummary, RetrieveOnly, EventMem, AnchorMem, RandomMem, and StructuredMem. AnchorMem has the highest endpoint score (mean r=0.005), is best across 5/5 voxel budgets and 5/5 ridge penalties, and wins the fold-level comparison against ShortWindow in 3/3 held-out stories. A matched RandomMem control is competitive at short prefixes but falls behind AnchorMem at the endpoint, across all ridge penalties, and across all voxel budgets. The supported claim is narrow but positive: under a matched context budget, memory selection changes the representation measured by this pilot, and sparse anchor memory is the most stable policy family tested here.