proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersHongyang Du, Yunfei Xie, Junjie Ye1 min readpaperadvanced

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Summary

FuseReg replaces the fixed heuristic of selecting encoder layers for representation autoencoders with a regularization that trains on random subsets of layers, making the downstream decoder robust to any fusion. This yields higher reconstruction quality (PSNR) and lowers unguided generation FID by up to 29% on ImageNet‑256, all without changing the pretrained visual encoder.

  • FuseReg trains decoders on random subsets of encoder layers, making them robust to any layer fusion without retraining.
  • A single FuseReg decoder on ImageNet‑256 (DINOv3‑L) achieves higher PSNR than decoders specialized to fixed fusions.
  • Replacing the decoder with FuseReg reduces unguided gFID by 27% for the RAEv2 DiT‑XL generator.
  • Joint regularization of encoder and diffusion stages cuts unguided gFID by 29% on DiT‑Base.

Practitioners building image generation or reconstruction systems with pretrained visual encoders should care because FuseReg improves both fidelity and generation quality without extra encoder training.

7/10

Related reading

  1. Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

    This paper introduces a training-adaptive Convolutional Sparse Coding (CSC) framework where the sparsity coefficient is learned end-to-end via FISTA unfolding. It uses an information bottleneck perspective to balance representation compression and content preservation, showing improved robustness to input perturbations on CIFAR and ImageNet.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene

    Mira-Scene proposes a compositional 3D scene reconstruction pipeline that replaces sparse pose regression with a dense Canonical Coordinate Map (CCM) linking image pixels to bounded object space. Coupled with a diffusion transformer, it achieves up to 40% higher 3D‑IoU than prior methods without scene‑level layout annotations.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

    The paper presents SemanTok, a flexible video tokenizer that injects frozen DINO features and reconstructs them from any token prefix, achieving strong semantic alignment and fidelity. A 201 M SemanTok AR model matches or exceeds a 3.4× larger VideoFlexTok baseline, with cheaper short‑prefix prediction and better generation quality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Deep Residual Learning for Image Recognition

    The paper proposes reformulating deep layers as residual functions with identity shortcut connections, making it easy to train networks far deeper than before. Using this design, a 152‑layer ResNet achieved 3.57% top‑5 error on ImageNet, winning ILSVRC 2015.

    Hall of Famearxiv.org42 minpaper