Hugging Face Daily PapersDonghao Zhou, Haoyang He, Fan Zhang1 min readpaperadvanced
ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing
Summary
ThinkV2V adds an explicit reasoning step using a multimodal LLM before conditioning a video diffusion model, trained with a progressive curriculum and refined at inference time. The 5B DiT model trained on a new 150K dataset outperforms larger baselines on complex instruction-guided edits.
- ThinkV2V inserts an MLLM "thinking" phase to generate refined conditioning signals for video diffusion editing.
- Progressive Curriculum Training gradually introduces reasoning complexity during model training.
- Inference-Time Thinking Scaling iteratively refines prompts and selects the most reliable edit.
- The authors release ThinkV2V-150K, a large instruction‑video dataset, and ThinkV2V-Bench for evaluating implicit and causal edits.
Engineers building instruction‑guided video editing tools should care because the paper shows how to harness LLM reasoning to handle implicit, causal edits that prior methods miss.
7/10