Hugging Face Daily PapersHongyang Du, Yunfei Xie, Junjie Ye1 min readpaperadvanced
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
Summary
FuseReg replaces the fixed heuristic of selecting encoder layers for representation autoencoders with a regularization that trains on random subsets of layers, making the downstream decoder robust to any fusion. This yields higher reconstruction quality (PSNR) and lowers unguided generation FID by up to 29% on ImageNet‑256, all without changing the pretrained visual encoder.
- FuseReg trains decoders on random subsets of encoder layers, making them robust to any layer fusion without retraining.
- A single FuseReg decoder on ImageNet‑256 (DINOv3‑L) achieves higher PSNR than decoders specialized to fixed fusions.
- Replacing the decoder with FuseReg reduces unguided gFID by 27% for the RAEv2 DiT‑XL generator.
- Joint regularization of encoder and diffusion stages cuts unguided gFID by 29% on DiT‑Base.
Practitioners building image generation or reconstruction systems with pretrained visual encoders should care because FuseReg improves both fidelity and generation quality without extra encoder training.
7/10
