1
A Zeroth-Order Paradigm for LLM Preference Alignment
The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth‑order method that uses comparison oracles to align LLMs without a differentiable loss. Experiments on several LLM families show it improves win rates and mitigates likelihood displacement compared to direct alignment approaches.
Hugging Face Daily Papersarxiv.org1 minpaper
