Gemini 2.5 Flash-Lite is now ready for scaled production use
Google DeepMind 9 months ago 25 ● 3 sources
Google released the stable version of Gemini 2.5 Flash-Lite, a lightweight model designed for fast, cost-efficient inference across tasks like translation and classification. The model costs $0.10 per million input tokens and $0.40 per million output tokens, with a 1 million-token context window and support for native tools including grounding with Google Search and code execution. Users can deploy the model by specifying "gemini-2.5-flash-lite" in their code, with Google retiring the preview alias on August 25th.