proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJingxuan Wu, Yuzhe Yang, Yiqiao Huang1 min readpaperadvanced

MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

Summary

MemFold learns a fixed‑size soft memory of K vectors for each user by optimizing the downstream task behavior via on‑policy rewards and a confidence‑gated teacher distillation. It improves accuracy on long‑context persona benchmarks and transfers to other evaluation sets without extra training.

  • Compresses query‑conditioned user history into K continuous vectors that serve as the model's memory interface.
  • Training combines group‑relative reward signals with on‑policy distillation from a frozen textual‑memory teacher, avoiding extra decoding at inference.
  • Achieves state‑of‑the‑art accuracy on PersonaMem‑32K/128K, with larger margins as history length grows.
  • Ablations show the reward term drives most gains; the teacher signal adds a smaller but consistent boost.

Engineers building long‑context LLM assistants need efficient, task‑aligned memory mechanisms to retain user preferences without exploding token length.

6/10

Related reading

  1. Do LLMs Have the Memory of a Goldfish?

    The article explains that LLMs don’t have persistent personal memory; all “memory” is supplied by the surrounding application via the context window, summaries, or external storage. It outlines the distinction between trained weights, working‑memory (token context), and persistent application memory, shows how to construct API calls to preserve conversation state, and discusses the cost and laten…

    ByteByteGobytebytego.com12 min
  2. MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

    The paper proposes Mixture of Memory Embeddings (MoME), a context‑aware sparse lookup that replaces each token’s single memory row with a gated mixture of multiple slots. Experiments on Llama‑3, MobileLLM and Qwen3 show MoME outperforms existing memory‑embedding baselines at equal parameter and FLOP budgets and exhibits interpretable routing for polysemous tokens.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper