proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJaihyun Lew, Mingi Jung, Minjun Park1 min readpaperadvanced

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

Summary

FoMo introduces a fully automated method for generating perceptual distance labels between image pairs, eliminating the need for human annotation. It leverages the "forking moment" in a diffusion model's generative trajectory, where early divergence indicates large perceptual differences and late divergence indicates subtle ones.

  • FoMo automates perceptual distance labeling using diffusion model dynamics, bypassing expensive human annotation.
  • The "forking moment" in a diffusion model's generation process correlates with human perception of image differences.
  • Early forking implies coarse structural differences, while late forking implies fine detail differences.
  • FoMo generates pointwise labels, providing richer training data than traditional pairwise comparisons.

This paper is crucial for ML engineers and researchers developing or evaluating generative AI models, as it offers a scalable and objective way to measure image quality and perceptual similarity without relying on costly and inconsistent human judgments.

8/10

Related reading

  1. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper