Hugging Face Daily PapersMario Sanz-Guerrero, Minh Duc Bui, Manuel Mager1 min readpaperadvanced
Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
Summary
The authors demonstrate that many LLM APIs silently embed the current date into system prompts, and this hidden variable alone shifts performance by up to 14% on math reasoning and alters leaderboard rankings day‑to‑day. Evaluation pipelines must fix or remove the date to ensure reproducible results.
- System prompts automatically include the current date, which users cannot control.
- Performance varies with the date: up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on MT.
- Model rankings on leaderboards can flip solely due to the date change.
- Chain‑of‑thought and few‑shot prompting do not mitigate the effect; CoT can even amplify it.
Anyone benchmarking or comparing LLMs needs stable, reproducible results, and hidden date drift directly undermines that.
8/10



