Hugging Face Daily PapersRui Liu, Bhavin Jawade, Haoqi Li1 min readpaperadvanced
Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing
Summary
Align Then Reason (ATR) is a multilingual lip-sync judge for dubbing that works without dubbed audio. It first aligns lip movements to phonetic units, then uses an LLM to reason about content and timing, achieving significant AUC improvements over baselines across seven languages and downstream tasks.
- Dubbing quality control requires a reference-free judge using only silent video and text, as dubbed audio may not exist.
- Existing visual speech recognizers and video-language models are insensitive to temporal errors in lip-sync.
- ATR employs a two-stage approach: monotonic alignment of frame-level lip representations to phonetic units, followed by LLM reasoning.
- An alignment scorer provides the LLM with both local and global evidence for joint content and timing judgment.
This method is crucial for automated, reference-free quality control in multilingual dubbing workflows, enabling efficient and accurate lip-sync evaluation before audio is even generated.
8/10