Explorerโ€บComputer Visionโ€บComputer Vision
Research PaperResearchia:202609.21007

MintAct: A Unified Visual Agent for Digital Environments

Mingfei Gao

Abstract

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds...

Submitted: September 21, 2026Subjects: Computer Vision; Computer Vision

Description / Details

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.


Source: arXiv:2609.22083v1 - http://arxiv.org/abs/2609.22083v1 PDF: https://arxiv.org/pdf/2609.22083v1 Original Link: http://arxiv.org/abs/2609.22083v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 21, 2026
Topic:
Computer Vision
Area:
Computer Vision
Comments:
0
Bookmark