Build with Nano Banana Pro, our Gemini 3 Pro Image model
Google DeepMind ● Covered by 2 sources
Google just launched Nano Banana Pro, its new Gemini 3 Pro Image model for developers. It reads text way better, checks facts via Google Search, and watermarks every image it makes.
Based on reporting by Google DeepMind — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google DeepMind is rolling out Nano Banana Pro, officially called Gemini 3 Pro Image, into paid preview for developers through the Gemini API in AI Studio and Vertex AI. It's the follow-up to Nano Banana (Gemini 2.5 Flash Image), which launched just a few months back and quickly became a favorite for things like character consistency and photo restoration. This time Google built the model on top of Gemini 3 Pro itself, and the jump in capability shows.
The headline feature is control. Developers can now dial in lighting, camera angle, focus, and color grading with a precision that older image models simply couldn't offer, and outputs come in 2K or 4K resolution for production-grade work. Google says the model can keep up to five people looking consistent across a set of images, blend six high-fidelity photos together, or combine as many as fourteen standard inputs into a single ad. That's a genuinely useful trick for anyone building marketing tools instead of just novelty generators.
Text rendering is the other big leap. Gemini 2.5 Flash Image struggled with legible, accurate text baked into images — most models do. Gemini 3 Pro Image handles logic and language well enough to produce clean text directly in generated scenes, and it understands the semantic context of an image well enough to translate signs, menus, or documents into another language while preserving the original layout and style. Google's food-word logo demo, where letters are built out of actual ingredients, is a decent showcase of how far this has come.
And because the model can ground its output in Google Search when needed, it's positioned as more than a picture generator — Google is pitching it for factual assets like biological diagrams or historical maps, where accuracy actually matters. There's a public infographic demo app that pulls this off in real time. Every image the model touches, generated or edited, gets a SynthID watermark, which is Google's way of getting ahead of the inevitable questions about provenance as this stuff floods the internet.
Google is also pushing the model outward fast: it's baked into Antigravity, the company's new agentic coding platform, so coding agents can spin up UI mockups or visual assets before a human even reviews the code. Adobe and Figma are integrating it too. The pricing tradeoff is straightforward — Flash Image stays around for cheap, fast jobs, while Pro Image costs more and runs slower but delivers the quality bump for anyone building tools that actually need it.
My take — AI-written commentary, not fact-checked reporting
The text-rendering fix and Search grounding are the real story here, not the resolution bump — image models being unable to spell has been the industry's most embarrassing open secret for years. Baking SynthID into every output by default is the right move, and I'd like to see every major lab treat watermarking as non-negotiable rather than a PR afterthought. Watch how fast Adobe and Figma integrations turn this into invisible infrastructure that most people using it will never even know is Google underneath.
Read more about this at: Google DeepMind