Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Google DeepMind ● Covered by 2 sources
Google released Gemma 4 12B, a multimodal model that processes images and audio directly without separate encoders, designed to run on laptops with 16GB of RAM. The model delivers performance comparable to Google's larger 26B model while requiring less than half the memory footprint. Developers can now build multimodal and agentic applications locally on consumer hardware without cloud dependencies.