1
MintAct: A Unified Visual Agent for Digital Environments
MintAct is a family of vision-language models (2B-8B) that unifies UI grounding, multi-step navigation, and visual tool use across mobile, desktop, and web environments. It achieves state-of-the-art performance by leveraging a scalable environment and an asynchronous reinforcement learning infrastructure.
Hugging Face Daily Papersarxiv.org1 minpaper
