proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersIvan Kobyzev, Abbas Ghaddar, Ali Nasiri-Sarvi1 min readpaperadvanced

Fractional State Space Transition for Long Sequence Modeling

Summary

The paper introduces FRAC, a state‑space model that uses fractional dynamics to achieve power‑law long memory instead of exponential forgetting. Approximating the fractional kernel with a log‑spaced exponential sum yields an efficient recurrent module that improves long‑context performance in large language models.

  • FRAC replaces exponential decay in SSMs with power‑law memory using fractional dynamics.
  • Implements the fractional kernel as a log‑spaced sum of exponentials for efficient recurrent computation.
  • Supports parallel training and prefill while keeping a bounded autoregressive state.
  • 1.3B‑parameter language model experiments show consistent gains on long‑context tasks and parity on short contexts.

ML engineers and researchers building long‑context LLMs should care because FRAC offers a practical way to extend memory without sacrificing efficiency.

6/10

Related reading

  1. The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

    The paper surveys recent attention variants in large language models, introducing a five‑dimensional framework (Memory Representation, Update, Access, Readout, Integration) to compare them. It shows that modern LLMs increasingly treat contextual memory as a coordinated, multi‑layer resource rather than a single attention operator.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaling Laws for Neural Language Models

    This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

    Hall of Famearxiv.org67 minpaper
  3. Attention Is All You Need

    The paper proposes the Transformer, a sequence‑to‑sequence model that relies solely on self‑attention, eliminating recurrence and convolutions. It achieves state‑of‑the‑art translation BLEU scores while training orders of magnitude faster.

    Hall of Famearxiv.org27 minpaper