Hugging Face Daily PapersJingxuan Wu, Yuzhe Yang, Yiqiao Huang1 min readpaperadvanced
MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
Summary
MemFold learns a fixed‑size soft memory of K vectors for each user by optimizing the downstream task behavior via on‑policy rewards and a confidence‑gated teacher distillation. It improves accuracy on long‑context persona benchmarks and transfers to other evaluation sets without extra training.
- Compresses query‑conditioned user history into K continuous vectors that serve as the model's memory interface.
- Training combines group‑relative reward signals with on‑policy distillation from a frozen textual‑memory teacher, avoiding extra decoding at inference.
- Achieves state‑of‑the‑art accuracy on PersonaMem‑32K/128K, with larger margins as history length grows.
- Ablations show the reward term drives most gains; the teacher signal adds a smaller but consistent boost.
Engineers building long‑context LLM assistants need efficient, task‑aligned memory mechanisms to retain user preferences without exploding token length.
6/10
