Gemini 2.5 Flash-Lite is now ready for scaled production use
Google DeepMind ● Covered by 3 sources
Google released the stable version of Gemini 2.5 Flash-Lite, a lightweight model designed for fast, cost-efficient inference across tasks like translation and classification. The model costs $0.10 per million input tokens and $0.40 per million output tokens, with a 1 million-token context window and support for native tools including grounding with Google Search and code execution. Users can deploy the model by specifying "gemini-2.5-flash-lite" in their code, with Google retiring the preview alias on August 25th.
Why it matters
Gemini 2.5 Flash-Lite, previously in preview, is now stable and generally available. This cost-efficient model provides high quality in a small size, and includes 2.5 family features like a 1 million-token context window and multimodality.