proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersRui Liu, Bhavin Jawade, Haoqi Li1 min readpaperadvanced

Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing

Summary

Align Then Reason (ATR) is a multilingual lip-sync judge for dubbing that works without dubbed audio. It first aligns lip movements to phonetic units, then uses an LLM to reason about content and timing, achieving significant AUC improvements over baselines across seven languages and downstream tasks.

  • Dubbing quality control requires a reference-free judge using only silent video and text, as dubbed audio may not exist.
  • Existing visual speech recognizers and video-language models are insensitive to temporal errors in lip-sync.
  • ATR employs a two-stage approach: monotonic alignment of frame-level lip representations to phonetic units, followed by LLM reasoning.
  • An alignment scorer provides the LLM with both local and global evidence for joint content and timing judgment.

This method is crucial for automated, reference-free quality control in multilingual dubbing workflows, enabling efficient and accurate lip-sync evaluation before audio is even generated.

8/10

Related reading

  1. onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

    onPanda is an interactive annotation tool that lets humans correct LLM outputs token‑by‑token, then resumes generation from the corrected prefix. In a controlled study it cut median annotation time by 52% and the authors release a token‑level correction dataset (Panda‑CVL) for on‑policy fine‑tuning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    OmniSeek is a framework that turns a large language model into a multi‑turn audio‑visual reasoning agent by letting it decide when to look or listen and which temporal windows to fetch as evidence. The authors train it on a synthetic 170K trajectory dataset, then refine with reinforcement learning and an Audio‑Visual Necessity loss, reporting consistent gains on several benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper