1
ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing
ThinkV2V adds an explicit reasoning step using a multimodal LLM before conditioning a video diffusion model, trained with a progressive curriculum and refined at inference time. The 5B DiT model trained on a new 150K dataset outperforms larger baselines on complex instruction-guided edits.
Hugging Face Daily Papersarxiv.org1 minpaper
