Microsoft Fara1.5-27B Enables Multimodal Browser Control via Screenshots
July 22, 2026
Fara1.5-27B uses vision-only perception to execute web tasks by predicting grounded actions like clicks and typing from screenshots. The model is fine-tuned from Qwen3.5-27B using synthetic data from the FaraGen1.5 multi-agent pipeline and is optimized for deployment with MagenticLite.
HOW THIS AFFECTS YOU
●
builderYou can build browser agents that operate via visual screenshots rather than relying on brittle DOM parsing.
●
researcherThis provides a framework for studying agentic behavior using synthetic, verified trajectory data.