Hugging Face Daily PapersJenna Russell, Ben Glickenhaus, Katherine Thai2 min readpaperadvanced
How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Summary
AI-generated web text now comprises a significant portion of pretraining data. This paper shows that adding "wild" AI tokens initially helps data-starved models but quickly becomes harmful, while immediately harming models with ample human data, necessitating a new scaling law.
- AI-generated text makes up over 30% of web data by August 2026, impacting pretraining.
- For data-starved models, AI text initially lowers loss on human text, but then reverses to harm.
- For models with sufficient human data, AI text immediately increases loss on human text.
- A new scaling law with separate benefit/harm terms better predicts AI text impact than Chinchilla.
This paper is crucial for anyone designing or training large language models, as it quantifies the complex and often detrimental impact of increasingly prevalent AI-generated data on model performance and provides actionable guidance.
8/10

