SyncAI.news, a Varaisys broadcasting
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
GD

Google DeepMind

· 1 min read

AI LabsGoogle DeepMind

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

Jun 03, 2026

|

Gemma 4 12B is designed to bring high-performance multimodal intelligence directly to your laptop, combining mobile-first efficiency with advanced reasoning.

Olivier Lacombe

Director of Product Management, Google Deepmind

Gus Martins

Product Manager, Google DeepMind

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Today, we are introducing Gemma 4 12B, our latest model designed to bring agentic multimodal intelligence directly to laptops. Bridging the gap between our edge-friendly E4B and our more advanced 26B Mixture of Experts (MoE), Gemma 4 12B packages powerful capabilities inside a reduced memory footprint. It is also our first mid-sized model to feature native audio inputs.

Thanks to the developer community, Gemma 4 models have now crossed 150 million downloads. You’ve built everything from wearable robotic arms for physical assistance to enterprise-grade AI security. We're excited to see what you build with this latest addition.

Here’s an overview of what makes Gemma 4 12B unique:

  • Novel unified architecture: No multimodal encoders. The vision and audio inputs flow directly into the LLM backbone.
  • Advanced reasoning: Benchmark performance nearing our 26B model, unlocking powerful multi-step reasoning and agentic workflows.
  • Laptop ready: Small enough to run locally with just 16GB of VRAM or unified memory.
  • Open and accessible: Released under an Apache 2.0 license with support across the developer ecosystem.
  • Drafter-ready: Gemma 4 12B comes equipped with Multi-Token Prediction (MTP) drafters to reduce latency.

Together, these features bring advanced multimodal capabilities to everyday hardware without sacrificing speed or reasoning. Let's now take a closer look at how Gemma 4 12B achieves this.

Run state-of-the-art agents locally

Experience a uniquely efficient, unified architecture

Here is how Gemma 4 12B processes multimodal inputs natively:

Original source

This story was published by Google DeepMind. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on deepmind.google

Similar News