MemFold:Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

Compact memory. Personalized answers.

Remembering a user is not the same as acting on what they said. MemFold compresses a user’s history into a fixed set of soft vectors and trains the reader on its own answers, so the memory is judged by the responses it supports.

Learning to use what we remember

An assistant that serves the same person over many conversations has to answer from what that person has revealed: which preferences still hold, which were revised, and which constraints apply now. Personalization is therefore a property of the answer, not of the memory store.

Retaining that information and acting on it are different problems. Keeping the history as text makes the reader’s input grow with every interaction. Compressing it into a fixed number of vectors bounds that input, but such memories are usually trained to reconstruct text or imitate reference answers, which are sequences the reader never produced itself. Neither objective sees the mistakes the reader makes once the text is gone.

MemFold judges a compact memory by the answers it supports. The reader generates responses from soft memory and learns from two signals on those same responses: task rewards that score the outcome, and token-level guidance from a frozen teacher that reads the corresponding textual memory.

From interaction history to a useful memory

A writer first extracts the query-relevant evidence from the history visible when the question arrives. A Perceiver-style compressor maps this textual memory into K continuous vectors. The reader answers from the query and these vectors alone, so its memory input stays the same size however long the history grows.

01User historyVisible before the query
02Textual memoryQuery-relevant evidence
03256 soft vectorsA fixed reader interface

Supervised initialization teaches the reader to consume this interface. On-policy training then improves how it uses the interface, with two signals computed on the same student-generated responses (Fig. 1):

  1. Reward the outcome. Group-relative policy optimization (GRPO) scores each sampled answer on the task and reinforces the better ones.
  2. Guide individual tokens. A frozen reader of the textual memory re-scores the student’s sampled tokens. A detached confidence gate upweights tokens that the text supports more strongly than the soft memory does, and bounds the teacher’s influence on any single token.
FIG. 1 One set of student rollouts, two training signals. The teacher and compressor remain frozen during on-policy optimization; only the student adapter is updated.

The teacher is never sampled from. It only scores tokens the student has already produced, so supervision stays on the student’s own distribution and adds no autoregressive decoding. The teacher is removed at inference. The memory budget bounds the reader’s input, but building the memory still requires reading the history once.

Evaluating long-context personalization

We evaluate MemFold with three Qwen backbones on PersonaMem-32K and PersonaMem-128K, then test direct transfer to PrefEval and LongMemEval. It achieves the highest accuracy for every backbone and benchmark we report, tying the strongest baselines on PrefEval with Qwen2.5-7B. The margin widens with history length: for all three backbones, the lead over the best baseline is larger on PersonaMem-128K than on 32K, although the reader receives the same number of memory vectors in both settings.

Accuracy (%) · Higher is better

Benchmark accuracy for Qwen2.5-3B-Instruct
MethodIn-domainOut-of-domain
PersonaMem
32K
PersonaMem
128K
PrefEvalLongMemEval
Full Text46.021.912.926.6
xRAG36.055.88.510.2
AutoCompressor32.030.5——
MemGen54.066.113.33.8
GRPO68.058.411.326.0
OPSD54.029.612.826.2
MemFold70.088.419.932.4

On PersonaMem-128K, MemFold reaches 88.4% accuracy, 22.3 percentage points above the strongest baseline in this comparison.

Transfer beyond the training dataset

For PrefEval, we evaluate checkpoints trained on PersonaMem-32K. For LongMemEval, we use checkpoints trained and selected on LoCoMo. Neither setting includes target-dataset training. PrefEval keeps the preference-modeling task but changes the data and evaluation format; LongMemEval tests broader memory abilities such as multi-session reasoning, temporal relations, and knowledge updates. Source-domain gains do not always carry over: on PrefEval, GRPO and OPSD can fall below the untrained Full Text baseline, while MemFold stays robust across both settings.

The accuracy does not come from spending more at inference. The weighted-average token-equivalent cost is lower than Full Text for all three backbones. The saving does not hold everywhere: on PrefEval, histories are too short for reader-side savings to offset memory construction, and a few other individual comparisons also use more tokens.

Learning from the reader’s own responses

With Qwen2.5-3B-Instruct on PersonaMem-32K, MemFold improves accuracy faster per optimizer update and reaches comparable accuracy with fewer student rollouts (Fig. 2). Both objectives reuse the same sampled responses, and the teacher scores tokens without generating any.

FIG. 2 Training on PersonaMem-32K with Qwen2.5-3B-Instruct. The curves compare complete recipes with different inputs and initializations. They measure optimizer updates and student rollouts, not wall-clock time.

Where does the improvement come from?

In the component ablation, GRPO accounts for most of the task gain. Adding OPD raises mean accuracy from 71.9 to 75.4 on PersonaMem-32K and from 86.9 to 87.9 on PersonaMem-128K. Reader initialization is essential: without it, mean accuracy falls to 47.9 and 46.2, because the reader must first learn to consume soft memory before on-policy training can help. A text-space control trained with the same objectives on textual memory reaches 69.3 and 85.5, below the soft-memory model, which suggests that fixed-budget compression is not the main bottleneck in this setup.

These ablation scores average 16 sampled responses at temperature 1.0, so they are not directly comparable with the greedy-decoding main results. The rollout curves also exclude the cost of initialization and teacher forward passes.

Examining the memory interface

We probe the interface in two ways: by varying its capacity and by changing the memory the reader receives. Accuracy peaks at K = 256 on both PersonaMem settings and is nearly flat from 128 upward (Fig. 3A), so only the smallest budget is clearly too small. Increasing K to 512 does not help under this training recipe, which was not retuned for each budget.

Replacing the matched memory with shuffled or null memory causes large accuracy drops (Fig. 3B), so the reader depends on the memory built for this instance rather than on task priors alone. The intervention shows that the memory is used. It does not show which facts the vectors preserve, or whether the reader picks the right version of a preference that changed over time.

FIG. 3A Memory budget and accuracy with Qwen3-4B. Both PersonaMem settings peak at K = 256 under the tested recipe.
FIG. 3B Memory interventions with Qwen3-4B. Matched memory outperforms shuffled and null memory on both PersonaMem settings.
FIG. 4 Token cost relative to K = 256 with Qwen3-4B. Labels show absolute counts in thousands. All budgets include the history-reading pass.

What does the memory budget cost?

Moving K from 64 to 512 barely changes end-to-end token-equivalent cost (Fig. 4); the variation is below one percent on PersonaMem-128K. Nearly all of the per-instance cost comes from reading the history once to build the textual memory, so enlarging the budget costs little and shrinking it saves little.

The textual-memory teacher itself scores well below the final student (Fig. 3B). That fits its role: it shows how the backbone reads uncompressed memory, but it is not an accuracy oracle. Its influence on each token is bounded, and GRPO drives most of the task-level gain.

What personalization changes in an answer

Aggregate scores do not show where personalization fails inside a response. In Fig. 6, the interaction history implies that the user dislikes wearable technology. Asked how to monitor fitness progress, GRPO and OPSD give relevant answers, yet both recommend wearables. MemFold instead suggests a journal, body measurements, and a progress photo log.

FIG. 6 A PrefEval example with Qwen3-4B. GRPO and OPSD recommend wearables, while MemFold offers alternatives consistent with the user’s preference.

The displayed preference summarizes the interaction history; the models never see it stated explicitly. This single example illustrates a failure pattern, not how often it occurs, and it does not establish an end-to-end efficiency gain.

Toward memory that supports better answers

MemFold judges a compact memory by the answers it supports rather than the text it reconstructs. Training the reader on its own responses, with task rewards and a frozen textual-memory teacher, gives the highest accuracy we measured across three backbones and both PersonaMem history lengths. The learned interface transfers to PrefEval and LongMemEval without target-domain training, and the reader still depends on instance-specific memory.

Open questions remain: which facts survive compression, how the reader resolves preferences that change over time, and how to reduce the cost of building the memory. Answering them will help turn compact memory into more reliable personalization.

Go deeper into MemFold.

Paper, code, and model resources.

BibTeX

@misc{wu2026memfold,
  title  = {{MemFold}: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization},
  author = {Jingxuan Wu and Yuzhe Yang and Yiqiao Huang and Chengzhi Liu and Qingni Wang and Chengxuan Qian and Shutong Wu and Jiawei Zhang and Xin Eric Wang},
  year   = {2026},
  url    = {https://MemFold.github.io/}
}