The Fastest and Most Affordable Gemini 2.5 Flash-Lite Is Here
On July 22, 2025, Logan Kilpatrick and Zach Gleicher announced on the Google Developers Blog that the Gemini 2.5 Flash-Lite model had reached a stable release and was now generally available. This model is the fastest and most affordable in the Gemini 2.5 family, designed specifically for latency- and cost-sensitive tasks such as translation or classification. Priced at just $0.10 per 1 million input tokens and $0.40 per 1 million output tokens, it can process large volumes of requests without high costs. In addition, the price of audio input has been reduced by 40% compared to the preview version.
Key Model Features
Gemini 2.5 Flash-Lite offers lower latency than the 2.0 Flash-Lite and 2.0 Flash models across a wide range of prompts. It provides a context window of up to 1 million tokens, enabling it to handle complex tasks. It supports native tools such as Grounding with Google Search, Code Execution, and URL Context. The model was trained on data up to January 2025, and advanced reasoning can optionally be enabled for more demanding use cases. It is multimodal, so it can process text, images, documents, video, and audio, with an input limit of 500 MB.
Real-World Use Cases
The model has already proven itself in practice since its preview release. Satlyt uses it for a decentralized space computing platform, achieving a 45% reduction in latency for on-orbit diagnostics and a 30% decrease in energy consumption. HeyGen uses it to automate video planning, optimize content, and translate into more than 180 languages for global personalization. DocsHound processes long videos and extracts thousands of screenshots with low latency, accelerating documentation creation. Evertune analyzes brand representation in AI models and quickly generates reports thanks to the model's high speed.
How to Get Started with the Model
To use it, specify “gemini-2.5-flash-lite” in your code. If you have been using the preview alias, switch to the stable version by August 25, 2025, when the preview will be removed. The model is available in Google AI Studio and Vertex AI. According to related online information, such as Google's technical report, it is ideal for high-throughput, low-cost production deployments, balancing speed, price, and quality.




