proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersMario Sanz-Guerrero, Minh Duc Bui, Manuel Mager1 min readpaperadvanced

Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

Summary

The authors demonstrate that many LLM APIs silently embed the current date into system prompts, and this hidden variable alone shifts performance by up to 14% on math reasoning and alters leaderboard rankings day‑to‑day. Evaluation pipelines must fix or remove the date to ensure reproducible results.

  • System prompts automatically include the current date, which users cannot control.
  • Performance varies with the date: up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on MT.
  • Model rankings on leaderboards can flip solely due to the date change.
  • Chain‑of‑thought and few‑shot prompting do not mitigate the effect; CoT can even amplify it.

Anyone benchmarking or comparing LLMs needs stable, reproducible results, and hidden date drift directly undermines that.

8/10

Related reading

  1. Learning to solve hard problems in RL for LLMs by never giving up

    The post introduces the *Matthew Effect* in RL‑fine‑tuning of LLMs—performance gains concentrate on tasks the model already solves— and proposes *Never Give Up* (NGU), an adaptive sampling scheme that uses a small k for easy prompts and retries hard prompts with a high‑probability “never give up” loop. Experiments on math (AIME, GSM8k), code (Manufactoria), and larger‑scale setups (DeepScaler) sh…

    Hacker News front pagegithub.io11 minHN1179
  2. Prompts aren’t Real

    The talk argues that prompt engineering is a dead‑end and proposes building large evaluation/optimization pipelines (pass^k testing, adversarial scenario generation, automated prompt optimization) to make LLM agents reliable. It describes a workflow: generate tests, run them with/without a new “skill”, feed results to a genetic optimizer that mutates prompts, validate on hold‑out tests, and itera…

    Hacker News front pageevaluation.club24 mintalkHN11757
  3. Last 3 days: AI Evals, October cohort

    This post announces a live, hands-on course on building reliable evaluation systems for production AI agents, with enrollment closing soon. The course covers designing evals for quality, safety, cost, and latency, and building LLM-as-a-Judge systems.

    ByteByteGobytebytego.com1 minrelease