1
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
The paper proposes Mixture of Memory Embeddings (MoME), a context‑aware sparse lookup that replaces each token’s single memory row with a gated mixture of multiple slots. Experiments on Llama‑3, MobileLLM and Qwen3 show MoME outperforms existing memory‑embedding baselines at equal parameter and FLOP budgets and exhibits interpretable routing for polysemous tokens.
Hugging Face Daily Papersarxiv.org1 minpaper
