Google DeepMind releases Gemma 4 12B multimodal model and DiffusionGemma experimental text generation model
Model release Provisional 95% confidence first seen
Google DeepMind released two new models: Gemma 4 12B, a unified multimodal model that processes images and audio without separate encoders and runs on consumer hardware with 16GB RAM, and DiffusionGemma, a 26B experimental text generation model that generates text blocks simultaneously for up to 4x faster inference speed on GPUs. Both models are designed for local deployment on consumer and modest server hardware.