proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJenna Russell, Ben Glickenhaus, Katherine Thai2 min readpaperadvanced

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Summary

AI-generated web text now comprises a significant portion of pretraining data. This paper shows that adding "wild" AI tokens initially helps data-starved models but quickly becomes harmful, while immediately harming models with ample human data, necessitating a new scaling law.

  • AI-generated text makes up over 30% of web data by August 2026, impacting pretraining.
  • For data-starved models, AI text initially lowers loss on human text, but then reverses to harm.
  • For models with sufficient human data, AI text immediately increases loss on human text.
  • A new scaling law with separate benefit/harm terms better predicts AI text impact than Chinchilla.

This paper is crucial for anyone designing or training large language models, as it quantifies the complex and often detrimental impact of increasingly prevalent AI-generated data on model performance and provides actionable guidance.

8/10

Related reading

  1. Training Compute-Optimal Large Language Models

    The paper derives a compute‑optimal scaling law showing model size and training tokens should grow together, and validates it by training a 70B‑parameter model (Chinchilla) on 1.4 T tokens that outperforms much larger LLMs.

    Hall of Famearxiv.org66 minpaper
  2. Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

    Vercel’s September AI Gateway Production Index shows open‑weight models processing 56% of token volume (up from 7% in Dec 2025) while accounting for only 14% of spend. Token price fell 23.2% month‑over‑month. Anthropic’s Opus 5 captured 22.5% of spend, overtaking Fable 5 which dropped to 4.9%. OpenAI’s new GPT‑6 Astra grabbed ~7.7% of total gateway spend in its first 12 days, more than double Ant…

    Vercelvercel.com6 minHN2
  3. Scaling Laws for Neural Language Models

    This paper empirically studies scaling laws for neural language model performance, finding that cross-entropy loss scales as a power-law with model size, dataset size, and compute. It shows that optimal compute-efficient training involves using very large models, training on relatively modest data, and stopping significantly before convergence.

    Hall of Famearxiv.org67 minpaper