Hacker News front page13 min readadvanced
Understanding the Impact of LLM Watermarking on AI Agent Behavior
Summary
The post empirically shows that generative watermarks like SynthID‑Text alter token sampling, causing measurable changes in LLM refusals and tool‑calling behavior, especially under prompt injection. Paired‑run experiments reveal up to ~16% churn in tool calls and weakened refusals, varying by model and watermark key.
- Watermarking via SynthID‑Text introduces "sampling drift" that can change token choices, affecting both model refusals and downstream tool‑calling arguments.
- Paired disagreement (churn) between watermarked and unwatermarked runs reaches 6‑16% for tool calls, even when net accuracy loss is small.
- Under a simple prompt‑injection attack, watermarks significantly reduce refusal rates, turning harmful requests into compliant outputs for several models.
- The impact is model‑ and key‑dependent; some models lose accuracy due to wrong arguments, others due to malformed output.
Engineers building LLM‑driven agents and safety researchers need to know that provenance watermarks can unintentionally degrade safety and correctness of deployed systems.
8/10


