proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersDonghao Zhou, Haoyang He, Fan Zhang1 min readpaperadvanced

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Summary

ThinkV2V adds an explicit reasoning step using a multimodal LLM before conditioning a video diffusion model, trained with a progressive curriculum and refined at inference time. The 5B DiT model trained on a new 150K dataset outperforms larger baselines on complex instruction-guided edits.

  • ThinkV2V inserts an MLLM "thinking" phase to generate refined conditioning signals for video diffusion editing.
  • Progressive Curriculum Training gradually introduces reasoning complexity during model training.
  • Inference-Time Thinking Scaling iteratively refines prompts and selects the most reliable edit.
  • The authors release ThinkV2V-150K, a large instruction‑video dataset, and ThinkV2V-Bench for evaluating implicit and causal edits.

Engineers building instruction‑guided video editing tools should care because the paper shows how to harness LLM reasoning to handle implicit, causal edits that prior methods miss.

7/10

Related reading

  1. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Reasoning with Image Generation

    ReImaGin proposes using image generation models as a flexible visual reasoning mechanism for multimodal LLMs, moving beyond rigid visual tools. It consistently outperforms text-only and specialist vision-tool baselines by up to 25% on diverse visual reasoning tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Streaming Video Editing with Easy Adaptation

    The paper introduces SVEET, a framework that adapts a pretrained bidirectional video diffusion model for streaming video editing via an auxiliary branch with temporally independent self‑attention and a decoupled orthogonal training scheme. It achieves real‑time 15 FPS editing on a single H100 GPU without extra acceleration.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

    Where‑OPD introduces on‑policy self‑distillation for multimodal LLMs where the teacher receives spatially grounded textual guidance from synthetic scenes. Post‑training on these annotation‑free scenes improves real‑world vision‑language benchmarks by an average of 3.23 points.

    Hugging Face Daily Papersarxiv.org1 minpaper