proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZhiyu Xu, Weilong Yan, Yufei Shi1 min readpaperadvanced

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Summary

The paper presents AV‑GRPO, a diffusion‑based reinforcement‑learning framework that treats audio and video generation as separate but coordinated tasks by anchoring rollouts to each modality and freezing the opposite tower during optimization. Evaluated on the new 5DAV dataset and benchmarks, it achieves higher fidelity, better text‑modality alignment, and tighter audio‑video sync than the previo…

  • Modality‑anchored rollouts decouple reward signals per modality, stabilizing training difficulty.
  • Trajectory‑locked frozen‑tower optimization cuts compute cost and improves credit assignment.
  • Adaptive objectives and perturbation strengths are tuned to each modality's dynamics.
  • The 5DAV dataset provides difficulty‑controlled, decoupled samples across five dimensions for systematic training.

Researchers and engineers building multimodal generative systems should care because the method improves cross‑modal quality while reducing training cost.

7/10

Related reading

  1. DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

    DACA‑GRPO adds denoising‑aware credit assignment to GRPO‑style RL trainers for diffusion LLMs. It computes per‑token importance scores from intermediate denoising steps and uses stratified masking to reduce mean‑field bias in likelihood estimates. Plug‑and‑play on three existing GRPO methods, it yields consistent gains on seven downstream tasks (up to +5.6 pp math, +7.4 pp code, +36.3 pp constrai…

    Apple Machine Learning Researchapple.com1 minpaper
  2. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

    OmniVChat defines native audio‑visual dialogue where a model consumes raw audio and video streams and replies in text. The authors build OmniVChat‑Studio, a multi‑agent simulator that generates single‑ and multi‑turn audio‑visual conversations, and use it to create OmniVChat‑Bench, a benchmark covering five dialogue abilities. They also propose OmniVChat‑RL, a reinforcement‑learning reward that b…

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Video Generation Models: A Survey of Post-Training and Alignment

    This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper