On February 12, Google unveiled a major update to Gemini 3 Deep Think, a specialized artificial intelligence mode designed to solve complex problems in science, research, and engineering. The model was developed in close collaboration with scientists and researchers to tackle challenging research problems that often lack clear rules or a single correct solution, and where data tends to be incomplete or inaccurate.
The updated Deep Think is now available in the Gemini app to Google AI Ultra subscribers. For the first time, the model is also being made available through the Gemini API to select researchers, engineers, and businesses, who can express interest in early access.
Academic benchmark results
The Gemini 3 Deep Think model achieves unprecedented results on the most challenging academic benchmarks. On Humanity's Last Exam, which is designed to test the limits of modern AI models, it achieved 48.4% without using tools. On the ARC-AGI-2 benchmark, it scored 84.6%, as confirmed by the ARC Prize Foundation.
In programming, the model achieved an Elo rating of 3455 on the Codeforces platform, an exceptional performance in competitive programming. At the 2025 International Mathematical Olympiad, it reached gold-medal level.
The model also excels across broad scientific fields such as chemistry and physics. It achieved gold-medal-level results on the written portions of the 2025 International Physics Olympiad and the 2025 International Chemistry Olympiad. In advanced theoretical physics, it scored 50.5% on the CMT-Benchmark.
Use in research
Lisa Carbone, a mathematician at Rutgers University, is working on the mathematical structures needed by the high-energy physics community to connect Einstein's theory of gravity with quantum mechanics. In a field with very little existing training data, she used Deep Think to review a highly technical mathematics paper. The model successfully identified a subtle logical error that had previously gone unnoticed during human peer review.
At Duke University, the Wang Lab used Deep Think to optimize manufacturing methods for complex crystal growth with the goal of potentially discovering semiconductor materials. Deep Think successfully proposed a recipe for growing thin films larger than 100 μm, meeting a precise target that previous methods had been unable to achieve.
Anupam Pathak, head of research and development in Google's Platforms and Devices division and former CEO of Liftware, tested the new Deep Think to accelerate the design of physical components. The model can turn a sketch into a 3D-printable reality – it analyzes the drawing, models the complex shape, and generates a file for creating the physical object using 3D printing.
Aletheia mathematical research agent
To solve mathematical research problems, the team created a research agent with the internal code name Aletheia, powered by the Gemini Deep Think model. The agent includes a natural-language verifier that identifies errors in candidate solutions and enables an iterative process of generating and revising solutions. The agent can acknowledge that it is unable to solve a problem, a key feature that improves efficiency for researchers.
The agent also uses Google Search and web browsing to navigate complex research, preventing false citations and computational inaccuracies when synthesizing published literature.
Since reaching the gold-medal standard at the IMO in July 2025, Gemini Deep Think has advanced rapidly and now achieves up to 90% on the advanced IMO-ProofBench test. Aletheia has already enabled several advances in mathematical research, including an autonomous research paper generated by AI without human intervention that calculates certain structural constants in arithmetic geometry known as eigenweights.
Breakthroughs in computer science and physics
The Gemini Deep Think model has also demonstrated its value in computer science and physics. While collaborating with experts on 18 research problems, the advanced model helped overcome long-standing obstacles in algorithms, machine learning, combinatorial optimization, information theory, and economics.
Key achievements include solving classic computer science problems such as Max-Cut and Steiner Tree using advanced tools from unrelated branches of continuous mathematics. The model also disproved a decade-old conjecture in online submodular optimization by constructing a highly specific combinatorial counterexample.
In cosmic string physics, the model found a new solution using Gegenbauer polynomials that naturally absorbed the singularities and collapsed an infinite series into a finite-sum closed form.
Human-AI collaboration
This work demonstrates that general-purpose foundation models using agentic workflows can serve as a powerful scientific partner. Under the guidance of expert mathematicians, physicists, and computer scientists, Gemini Deep Think demonstrates its usefulness in fields where complex mathematics, logic, and reasoning are essential.
The model acts as a force multiplier for human intellect, handling knowledge retrieval and rigorous verification so that scientists can focus on conceptual depth and creative direction. Whether refining proofs, finding counterexamples, or connecting separate fields, AI is becoming a valuable collaborator in the next chapter of scientific progress.



