proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersBowen Yang, Jingbo Zhou, Qinghong Miao1 min readpaperadvanced

FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models

Summary

FactorEngram proposes a factorized n-gram memory with basis-level contextual gating for LLMs. This design allows polysemous patterns to selectively retrieve relevant memory components from a shared dictionary, leading to improved language modeling and downstream task performance.

  • FactorEngram uses factorized n-gram memory, retrieving sparsity-regularized coefficients over a shared dictionary of basis vectors.
  • Basis-level contextual gating allows the backbone hidden state to modulate each memory component individually before reconstruction.
  • This design enables polysemous patterns to selectively use memory components, addressing limitations of prior monolithic lookup-based memory.
  • Parameter sharing occurs through a semantically relevant dictionary of basis vectors, not just hash collisions.

This paper offers a concrete, novel approach for scaling LLMs more efficiently by improving how external memory interacts with model context, relevant for researchers and practitioners.

8/10

Related reading

  1. Do LLMs Have the Memory of a Goldfish?

    The article explains that LLMs don’t have persistent personal memory; all “memory” is supplied by the surrounding application via the context window, summaries, or external storage. It outlines the distinction between trained weights, working‑memory (token context), and persistent application memory, shows how to construct API calls to preserve conversation state, and discusses the cost and laten…

    ByteByteGobytebytego.com12 min
  2. The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

    The paper surveys recent attention variants in large language models, introducing a five‑dimensional framework (Memory Representation, Update, Access, Readout, Integration) to compare them. It shows that modern LLMs increasingly treat contextual memory as a coordinated, multi‑layer resource rather than a single attention operator.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

    The paper proposes Mixture of Memory Embeddings (MoME), a context‑aware sparse lookup that replaces each token’s single memory row with a gated mixture of multiple slots. Experiments on Llama‑3, MobileLLM and Qwen3 show MoME outperforms existing memory‑embedding baselines at equal parameter and FLOP budgets and exhibits interpretable routing for polysemous tokens.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Scaling Laws for Neural Language Models

    This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

    Hall of Famearxiv.org67 minpaper
  5. Language Models are Few-Shot Learners

    This paper introduces GPT-3, a 175-billion-parameter autoregressive language model, demonstrating that scaling model size significantly improves few-shot learning. It achieves strong performance on many NLP tasks by conditioning on text instructions and examples, often without needing gradient updates or fine-tuning.

    Hall of Famearxiv.org166 minpaper