Hugging Face Daily PapersZhiyu Xu, Weilong Yan, Yufei Shi1 min readpaperadvanced
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
Summary
The paper presents AV‑GRPO, a diffusion‑based reinforcement‑learning framework that treats audio and video generation as separate but coordinated tasks by anchoring rollouts to each modality and freezing the opposite tower during optimization. Evaluated on the new 5DAV dataset and benchmarks, it achieves higher fidelity, better text‑modality alignment, and tighter audio‑video sync than the previo…
- Modality‑anchored rollouts decouple reward signals per modality, stabilizing training difficulty.
- Trajectory‑locked frozen‑tower optimization cuts compute cost and improves credit assignment.
- Adaptive objectives and perturbation strengths are tuned to each modality's dynamics.
- The 5DAV dataset provides difficulty‑controlled, decoupled samples across five dimensions for systematic training.
Researchers and engineers building multimodal generative systems should care because the method improves cross‑modal quality while reducing training cost.
7/10
