1
RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
RelateAnything is a 53 M‑parameter model that predicts arbitrary textual relations between any supplied image regions, using a text‑embedding bank instead of a fixed classifier. Trained on a new 4.3 M‑relation dataset, it outperforms prior open‑vocab methods by 2.3‑3.5× mean recall while running in ~20 ms per frame.
Hugging Face Daily Papersarxiv.org2 minpaper
