New Delhi: Google has introduced the full launch of its latest on-device AI model, Gemma 3n, which was first introduced in May 2025. The AI model brings advanced multimodal capabilities, including audio, image, video, and text processing, to smartphones and edge devices with limited memory and an internet connection. With the release, developers can now deploy AI features that previously required robust cloud infrastructure directly on phones and low-power devices.
The Gemma 3n features the latest architecture, known as MatFormer, an abbreviation for Matryoshka Transformer. Google explains that much like Russian nesting dolls, the model includes smaller, fully functional sub-models inside larger ones. This design enables developers to scale performance according to the available hardware. For example, Gemma 3n is available in two versions: E2B, which operates on as little as 2GB of memory, and E4B, which requires approximately 3 GB.
We’re fully releasing Gemma 3n, which brings powerful multimodal AI capabilities to edge devices. 🛠️
Here’s a snapshot of its innovations 🧵 pic.twitter.com/ARo2nHdUzC
— Google DeepMind (@GoogleDeepMind) June 26, 2025
Gemma 3n also introduces KV cache Sharing, which significantly speeds up the processing of long audio and video inputs. Google says this improves response times by up to two times, making real-time applications like voice assistants or video analysis much faster and more practical on mobile devices.
For the speech-based features, Gemma 3n utilises a built-in audio encoder adapted from Google’s Universal Speech Model. This enables it to perform tasks such as speech-to-text and language translation directly on a phone. Early tests have shown powerful results when translating between English and European languages, such as Spanish, French, Italian, and Portuguese.
The visual side of Gemma 3n is powered by MobileNet-V5, Google’s latest lightweight vision encoder. This system can handle video streams up to 60 frames per second on devices like the Google Pixel, enabling smooth real-time video analysis. Despite being smaller and faster, it outperforms previous vision models in both speed and accuracy.
Developers can access Gemma 3n through popular tools such as Hugging Face Transformers, Ollama, MLX, and LLama.cpp, among others. Google has also launched the Gemma 3n Impact Challenge, inviting developers to create applications that utilise the model’s offline capabilities. Winners will share a $150,000 prize pool.
Importantly, the model can operate entirely offline, meaning it doesn’t need an internet connection to work. This opens the door for AI-powered applications in remote areas or privacy-sensitive situations where cloud-based models aren’t viable. With support for over 140 languages and the ability to understand content in 35, Gemma 3n sets the latest standard for efficient, accessible on-device AI.









