Related reading
FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
FLAT is a pre‑training framework that encodes images and text into a shared 1D token sequence with variable length via nested dropout, enabling the same embeddings for cross‑modal retrieval and generation. It attains state‑of‑the‑art scores on COCO and Flickr30K for captioning, retrieval, and text‑to‑image generation, and supports interpolation and zero‑shot composed retrieval.
Hugging Face Daily Papersarxiv.org1 minpaperRetention-Constrained Post-Training Quantization of Cellpose-SAM for Stem Cell Microscopy
The paper proposes a retention‑constrained post‑training quantization benchmark for Cellpose‑SAM, showing that weight‑only W8A16 and a mixed W4/W8 scheme keep instance F1 scores while cutting model size 6.76×, whereas ternary quantization fails on most images.
Hugging Face Daily Papersarxiv.org1 minpaperTraining-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
This paper introduces a training-adaptive Convolutional Sparse Coding (CSC) framework where the sparsity coefficient is learned end-to-end via FISTA unfolding. It uses an information bottleneck perspective to balance representation compression and content preservation, showing improved robustness to input perturbations on CIFAR and ImageNet.
Hugging Face Daily Papersarxiv.org1 minpaperBreaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.
Hugging Face Daily Papersarxiv.org1 minpaper
