1
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
Ovis-Embedding is a unified omni‑modal embedding model that shares a single backbone for text, image, video, and audio, trained with contrastive loss, homogeneous‑source sampling, and embedding distillation. It achieves state‑of‑the‑art results on several multimodal retrieval benchmarks while offering low‑rank, dimension‑flexible embeddings.
Hugging Face Daily Papersarxiv.org1 minpaper
