proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameTomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean201327 min readpaperadvanced

Efficient Estimation of Word Representations in Vector Space

Summary

This paper introduces two novel log-linear model architectures for efficiently computing continuous word vector representations from very large datasets. These models achieve state-of-the-art accuracy on syntactic and semantic word similarity tasks with significantly lower computational cost than previous neural network approaches.

  • New log-linear models enable learning high-quality word vectors from billions of words in under a day.
  • The learned vectors capture linear regularities, allowing vector arithmetic for syntactic and semantic analogies.
  • Computational complexity is minimized by avoiding non-linear hidden layers and using hierarchical softmax with Huffman trees.
  • Parallel training on distributed frameworks like DistBelief is essential for scaling to massive datasets.

This foundational paper introduced the core ideas behind Word2Vec, revolutionizing how word embeddings are learned and applied in NLP by making them efficient and high-quality.

9/10

Related reading

  1. Scaling Laws for Neural Language Models

    This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

    Hall of Famearxiv.org67 minpaper
  2. The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

    The paper surveys recent attention variants in large language models, introducing a five‑dimensional framework (Memory Representation, Update, Access, Readout, Integration) to compare them. It shows that modern LLMs increasingly treat contextual memory as a coordinated, multi‑layer resource rather than a single attention operator.

    Hugging Face Daily Papersarxiv.org1 minpaper