From games to reality: Google DeepMind SIMA 2 learns inside 3D worlds

From games to reality: Google DeepMind SIMA 2 learns inside 3D worlds

New Delhi: Google’s DeepMind has recently unveiled a new SIMA 2, which is the latest version of its Scalable Instructable Multiworld Agent that can reason, collaborate with users, and learn autonomously inside the 3D virtual environments. The researchers described the release as a milestone in creating general and helpful AI agents. SIMA 2 incorporates the Gemini model as its core, which enables the agent to interpret instructions, understand high-level goals, and describe its planned actions. SIMA 2 can do more than respond to instructions; it can think and reason about them. The earlier version, SIMA 1, had been trained to execute more than 600 basic skills across various commercial games.


Training for the SIMA 2 used both human demonstrations and labels generated by Gemini. This approach allows the agent to explain what it intends to do and how it plans to complete the task. Interactions now feel less like giving commands and more like collaborating with the companion who can reason about the task at hand. Testing showed improved generalisation, with the SIMA 2 carrying out complex instructions and succeeding in games it had never encountered, including the Viking survival title ASKA and the research environment MineDojo.

This agent could also apply concepts learned in one game, such as mining, to comparable actions in other environments. Researchers noted that SIMA 2 has reduced much of the performance gap between AI and human players across the evaluation tasks. SIMA 2 was combined with the Genie 3, a model that creates the latest 3D models from a single image or text prompt. The agent was able to orient itself and follow user instructions inside these automatically generated environments. A key capability in the new system is self-improvement. After the initial training on human demonstrations, SIMA 2 can shift to self-directed learning, using tasks and reward estimates generated by Gemini.

Google DeepMind noted that remaining limitations, including the difficulty with the very long, multi-step tasks, short interaction memory, and precision challenges when controlling games through virtual keyboard and mouse inputs. Visual understanding of complex 3D scenes also remains an area for improvement. SIMA 2 is a limited research preview for a small group of academics and game developers. Researchers stated the work may eventually inform robotics, where skills such as navigation, tool use, and collaborative task execution are essential.

Punit Panchal
Senior Editor

I’m a content writer specializing in tech, creating clear, engaging, and SEO-friendly content that simplifies complex topics. From emerging technologies to product insights, I focus on delivering value-driven content that connects with readers and ranks effectively.

Comments are closed