proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersAntoine Lorentz, Stéphane May, Valentine Bellet1 min readpaperadvanced

Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning

Summary

The authors adapt Stable Diffusion 3 with a pruned text stream and patch‑wise normalization to condition on photogrammetric DSMs and Pléiades imagery, refining elevation maps. In French city tests the method halves RMSE, reaching 3.45 m in‑context and 2.77 m on a held‑out city.

  • Modified Stable Diffusion 3 with a pruned text stream and patch‑wise normalization enables stable training on LiDAR elevation data.
  • Multimodal conditioning on photogrammetric DSMs and high‑resolution satellite imagery improves height prediction accuracy.
  • RMSE on dense urban areas drops from 6.00 m to 3.45 m in cities seen during training and from 4.16 m to 2.77 m on the unseen city of Bordeaux.
  • The technique transfers knowledge from natural‑image diffusion models to elevation maps without training a model from scratch.

Geospatial engineers and remote‑sensing teams can boost the quality of cheap photogrammetric DSMs using pretrained diffusion models, reducing reliance on expensive LiDAR surveys.

8/10

Related reading

  1. Denoising Diffusion Probabilistic Models

    This paper demonstrates that Denoising Diffusion Probabilistic Models (DDPMs) can generate high-quality images, achieving state-of-the-art FID scores on CIFAR10. It establishes a novel connection between DDPMs and denoising score matching, leading to a simplified training objective that predicts the noise added at each step.

    Hall of Famearxiv.org38 minpaper
  2. Adversarial Training for Pixel Diffusion

    Pixel diffusion models often underrepresent fine-scale image statistics. This paper demonstrates that adversarial post-training, by adding an adversarial loss to non-high-noise timesteps, effectively corrects this deficiency. It restores missing high-frequency content, jointly improving distribution fidelity, coverage, prompt alignment, and perceptual quality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

    FoMo introduces a fully automated method for generating perceptual distance labels between image pairs, eliminating the need for human annotation. It leverages the "forking moment" in a diffusion model's generative trajectory, where early divergence indicates large perceptual differences and late divergence indicates subtle ones.

    Hugging Face Daily Papersarxiv.org1 minpaper