MintAct Unified Vision-Language Models for Digital Agents
September 17, 2026
MintAct delivers 2B, 4B, and 8B scale models that unify UI grounding, multi-step navigation, and visual tool use across mobile, desktop, and web. The models match per-domain specialist performance through a scalable asynchronous RL infrastructure.
HOW THIS AFFECTS YOU
●
builderYou can deploy smaller, unified models that handle diverse digital tasks instead of managing multiple specialized agents.
●
founderThis reduces the complexity of building cross-platform automation products.