Microsoft’s Fara-7B sets stage for fully automated, AI-driven PC experience

Microsoft’s Fara-7B sets stage for fully automated, AI-driven PC experience

New Delhi: Microsoft has recently unveiled its Fara-7B, its first small language model, which is especially built to operate a computer the way a person does. The company has claims that the 7-billion parameter model matches or beats larger agentic systems on live web tasks while running locally with lower latency and stronger privacy. Fara-7B reads a webpage visually and completes tasks by clicking, typing and scrolling on predicted coordinates. It does not rely on accessibility trees or separate parsing layers. Microsoft has also stated that the model finishes tasks in about 16 steps on average, which is far fewer than many comparable systems. The model is trained on 145,000 synthetic trajectories generated through the Magentic-One framework and is built on Qwen2.5-VL-7B with the supervised fine-tuning.

The company has positions Fara-7B as an everyday computer use agent that can search, summarise, fill forms, manage accounts, book tickets, shop online, compare prices and find jobs or real estate listings. Microsoft is also releasing WebTailBench, the latest test set with 609 real-world tasks across 11 categories. Fara-7B leads all computer use models across every segment, including shopping, flights, hotels, restaurants, and multi-step comparison tasks.

The company offers two ways to run the model: Azure Foundry hosting lets users deploy Fara-7B without downloading weights or using their own GPUs. Advanced users can self-host through VLLM on GPU hardware. The evaluation stack relies on Playwright and an abstract agent interface that can plug in any model. Microsoft warns that Fara-7B is an experimental release and should be run in sandboxed settings without sensitive data. Microsoft launched Phi-4 multimodal and Phi-4 mini, the latest additions to its Phi family of small language models. Google DeepMind released the Gemini 2.5 Computer Use model, a specialised version of its Gemini 2.5 Pro AI that can interact with user interfaces. The model is available in preview through the Gemini API through Google AI Studio and Vertex AI Studio.

Punit Panchal
Senior Editor

I’m a content writer specializing in tech, creating clear, engaging, and SEO-friendly content that simplifies complex topics. From emerging technologies to product insights, I focus on delivering value-driven content that connects with readers and ranks effectively.

Comments are closed