Google Officially Launches Stable Versions of Gemini 2.5 Pro and Flash
Google has officially announced that its Gemini 2.5 Pro and Flash models are now generally available as stable versions, designed for production deployment. This family of models was designed as hybrid reasoning systems that deliver excellent performance while maintaining an optimal balance of cost and speed. The Gemini 2.5 Pro model remains unchanged from the May 6 version, while Gemini 2.5 Flash is identical to the May 20 version introduced at the Google I/O conference. Developers such as Spline and Rooms, along with organizations such as Snap and SmartBear, have already been using these latest versions in production environments for several weeks.
The stable versions give developers the confidence to build production applications. Both Gemini 2.5 Pro and Flash are now available not only through Google AI Studio and Vertex AI, but also in the Gemini app. In addition, Google has implemented custom versions of the 2.5 Flash-Lite and Flash models directly into its search engine, expanding their availability and use across various Google platforms.
Introducing the New Gemini 2.5 Flash-Lite Model
The latest addition to the Gemini 2.5 family is the Flash-Lite model, which is now available in preview. This model is the most cost-effective and fastest option in the entire Gemini 2.5 family. Flash-Lite was designed as a cost-effective upgrade to the previous 1.5 and 2.0 Flash models and offers better performance across most benchmarks while achieving lower time-to-first-token latency and faster token decoding speeds.
The 2.5 Flash-Lite model delivers higher overall quality than 2.0 Flash-Lite in programming, mathematics, science, reasoning, and multimodal benchmarks. It particularly excels at high-volume, latency-sensitive tasks such as translation and classification, where it achieves lower latency than 2.0 Flash-Lite and 2.0 Flash across a broad sample of prompts. The model includes the same features that make the entire Gemini 2.5 family useful, including the ability to enable "thinking" with different budgets, connect to tools such as Google Search and code execution, accept multimodal input, and support a context length of one million tokens.

Advanced "Thinking" Models with Dynamic Control
All Gemini 2.5 models are characterized as "thinking" models capable of working through their reasoning before responding, resulting in better performance and greater accuracy. Each model provides control over its "thinking" budget, allowing developers to choose when and how much the model should "think" before generating a response. This feature represents a significant innovation in artificial intelligence because it makes it possible to adapt the model's behavior to the specific needs of an application.
The Gemini 2.5 Flash-Lite model is optimized for cost and speed, so its "thinking" feature is disabled by default, unlike the other models. Nevertheless, it supports dynamic control of the thinking budget through an API parameter. Flash-Lite also supports all native Google tools, including Grounding with Google Search, Code Execution, and URL Context, in addition to function calling.

Pricing Changes and Pricing Structure
Google has made significant adjustments to the pricing of the Gemini 2.5 Flash model. Over the past year, the company's research teams have continued to push the Pareto frontier for the Flash model series. When the 2.5 Flash model was originally announced, Google had not yet finalized the capabilities of 2.5 Flash-Lite and launched it with separate pricing for "thinking" and "non-thinking," which caused confusion among developers.
With the launch of the stable version of Gemini 2.5 Flash, Google updated the pricing to $0.30 per 1 million input tokens (up from $0.15) and $2.50 per 1 million output tokens (down from $3.50). The company removed the price difference between "thinking" and "non-thinking" modes and retained a single pricing tier regardless of the number of input tokens. Although Google aims to maintain consistent pricing between preview and stable versions, this particular adjustment reflects the exceptional value of the Flash model, which continues to offer the best price-to-intelligence ratio on the market.
Continued Growth and Popularity of Gemini 2.5 Pro
Growth and demand for the Gemini 2.5 Pro model continue to follow the steepest curve of any model Google has ever released. To enable more customers to build production applications with this model, Google is making the June 5 version available as a stable release with the same optimal pricing structure as before. The Pro model is expected to excel in cases requiring the highest intelligence and greatest capabilities, such as programming and agentic tasks.
Gemini 2.5 Pro is at the core of many of the most popular developer tools, including Cursor, Bolt, Cline, Cognition, Windsurf, GitHub, Lovable, Replit, and Zed Industries. These tools leverage the advanced capabilities of the Pro model to provide high-quality services to their users. Google has also stated that it will share more information in the near future about scaling beyond the Pro model, indicating further developments in this area.
For users of older preview versions of the models, Google has set clear migration deadlines. Users of 2.5 Pro Preview 05-06 will have access to the model until June 19, 2025, after which it will be shut down. Users of 2.5 Pro Preview 06-05 can simply update the model string to "gemini-2.5-pro". Similarly, users of Gemini 2.5 Flash Preview 04-17 will retain the existing preview pricing until its scheduled discontinuation on July 15, 2025, when this endpoint will be shut down.



