proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersShaohua Dong, Zexuan Meng, Haiyan Sun1 min readpaperadvanced

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

Summary

Researchers introduce RGBD20K, a new large-scale dataset for RGB-D semantic segmentation with 20,000 image pairs and 160 fine-grained categories, featuring high-fidelity annotations. They also propose a novel score-purified fusion (SPF) method that achieves state-of-the-art performance on evaluated benchmarks.

  • RGBD20K offers 160 fine-grained categories, significantly more than prior RGB-D datasets like NYUv2 or SUN RGB-D.
  • The dataset contains 20,000 RGB-D image pairs, providing a substantially larger resource for training deep models.
  • Annotations in RGBD20K are high-fidelity, resulting from rigorous re-evaluation and correction of existing label noise.
  • A new score-purified fusion (SPF) method is introduced, achieving state-of-the-art results across evaluated benchmarks.

ML researchers and practitioners working on computer vision, especially 3D scene understanding and robotics, should care as this dataset and method can advance the development of more generalizable segmentation models.

7/10

Related reading

  1. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

    This post details how Google optimized spatio-temporal attention for video diffusion models on TPUs, turning theoretical sparsity into actual inference speedups. Key optimizations include specializing tile execution paths and carefully tuning tile sizes, resulting in significant latency reductions compared to dense attention.

    Google Developersgoogleblog.com11 min
  3. Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

    This paper evaluates GPT-6 Astra and five other frontier general-purpose AI systems across 34 computer vision capabilities and 55 benchmarks. It finds these systems excel at semantic interpretation and reasoning, but struggle with metric geometric accuracy, faithful reconstruction, and fine-grained specialized knowledge.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Denoising Diffusion Probabilistic Models

    This paper demonstrates that Denoising Diffusion Probabilistic Models (DDPMs) can generate high-quality images, achieving state-of-the-art FID scores on CIFAR10. It establishes a novel connection between DDPMs and denoising score matching, leading to a simplified training objective that predicts the noise added at each step.

    Hall of Famearxiv.org38 minpaper
  5. TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

    TAPe+ML v3 is a compact multi‑task vision system that replaces raw‑pixel processing with a structured TAPe representation. Using <100 k parameters, it achieves 84.7 mAP50 (65.3 mAP50‑95) on COCO detection, 80.7 mask mAP50 (58.4 mask mAP50‑95) on COCO segmentation, 92 % top‑1 on Imagenette and 89.9 % on ImageNet‑Real, while also showing robustness to distribution shift in video scene detection.

    Hugging Face Daily Papersarxiv.org1 minpaper