proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersShuai Yang, Luozhou Wang, Wei Huang1 min readpaperadvanced

LongLive-Plug: Once-for-All Distillation for Video Generation

Summary

LongLive-Plug introduces a once-for-all distillation framework for video generation models. It learns reusable capabilities as LoRAs on a base model, enabling training-free, plug-and-play deployment to 54 diverse downstream models.

  • Video model distillation is often repeated; LongLive-Plug distills capabilities once per backbone family.
  • It uses LoRAs to learn reusable capabilities like single-pass CFG, few-step sampling, and long-context error correction.
  • These LoRAs are plug-and-play, working on downstream models without retraining, even with added conditioning.
  • Verified on 54 downstream models across 3 backbone families and 8 task categories.

This framework significantly reduces the training overhead for specializing video diffusion models, benefiting researchers and engineers developing diverse video generation applications.

8/10

Related reading

  1. Video Generation Models: A Survey of Post-Training and Alignment

    This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Automating coherent long-form video generation

    Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

    Google Researchresearch.google10 min
  3. Mitigating the Length-Scaling Tax with Online Distillation

    The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. LOCI: Spatial Linear Memory for Streaming World Models

    LOCI is a hybrid spatial-memory architecture for streaming video world models, combining key-value caches with recurrent linear attention. It leverages projective camera geometry to condition memory operations, enabling more faithful reproduction of revisited content and significantly reducing peak memory usage for long videos.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper