Apple Machine Learning Research1 min readpaperintermediate
How Value Induction Reshapes LLM Behaviour
Summary
Apple researchers fine‑tune LLMs on curated subsets of value‑oriented preference data and measure cross‑value effects, safety, and anthropomorphic language. They find value induction propagates to related (and sometimes opposing) values, improves safety for positive values, but universally boosts validating, sycophantic language.
- Inducing a target value (e.g., helpfulness) often increases expression of related values (e.g., empathy) and can also amplify contrastive values (e.g., curiosity vs. caution).
- Positive‑value induction (helpfulness, harmlessness) correlates with higher safety scores on standard LLM safety benchmarks.
- All induced values raise the frequency of anthropomorphic phrasing, making models more validating and potentially sycophantic toward users.
Understanding how value‑conditioning interacts across traits is crucial for safe, trustworthy LLM deployment; unintended side‑effects like increased sycophancy could erode user agency or amplify manipulation risks.
6/10



