On February 11, 2026, Google DeepMind published a major update on Gemini 3 Deep Think. This specialized reasoning mode is no longer aimed at solving olympiad problems faster; it is built to work on open-ended research questions in mathematics, physics, and computer science, under the direction of professional researchers. The post was led by Thang Luong and Vahab Mirrokni, sponsored by Demis Hassabis and Jeff Dean, with Terence Tao among the thanked external experts. The byline list alone signals that DeepMind treats “AI for Science” as a frontline priority, not a side project.
From Olympiad Gold to Research-Grade Reasoning
Deep Think is Gemini 3’s dedicated reasoning mode. DeepMind says it reached gold-medal standard at the International Mathematical Olympiad in summer 2025, and that a later version matched that level at the International Collegiate Programming Contest. The star of this update is the January 2026 version: in internal evaluations it significantly outperformed the IMO-gold version while using less inference-time compute.
In other words, capability is still climbing while the cost curve bends down. For products built on reasoning modes, that matters more than any single leaderboard score.
Aletheia: A Math Agent That Admits Failure
The most interesting piece for engineers is Aletheia, a mathematical research agent built on Deep Think. It checks its own proofs with a natural-language verifier, uses Google Search and web browsing for literature work, and is explicitly designed to be able to admit failure.
That last property is critical to agent engineering. A research agent that cannot say “I could not do this” is more dangerous than one that fluently produces wrong proofs. A natural-language verifier also lets mathematicians audit machine reasoning in the format they already work in, instead of forcing them through formal proof systems.
90% on IMO-ProofBench Advanced and Reasoning That Scales
On IMO-ProofBench Advanced, Deep Think reaches up to 90% as inference-time compute increases, and DeepMind reports the scaling law holds through PhD-level exercises on its internal FutureMath Basic benchmark. The product implication is direct: when quality can be bought with compute, accuracy, latency, and cost become one set of tunable dials. You can segment pricing by scenario instead of shipping a single fixed model property.
The 18-Problem Ledger and Deliberate Underclaiming
DeepMind published an unusually concrete ledger: across 18 research problems, the work produced at least one ICLR 2026 acceptance. Highlights include Feng26, a fully autonomous paper computing eigenweight structure constants in arithmetic geometry; autonomous solutions to 4 of 700 open problems in Bloom’s Erdős conjecture database (including Erdős-1051, later generalized in BKKKZ26); and a three-item counterexample refuting a 2015 conjecture in online submodular optimization. The team also deployed it to assist STOC 2026 paper reviews.
Equally notable is the underclaiming: the report states explicitly that results so far are Level 2 “publishable quality,” with no Level 3 or Level 4 breakthroughs. In a hype-driven release culture, that restraint makes the whole report more credible, not less.
Availability and Limits
Deep Think is currently available in the Gemini app for Google AI Ultra subscribers, as Google confirmed through its official channels; on the API side it is limited to selected researchers, engineers, and enterprises through the Early Access Program, and has not entered the general pay-as-you-go Gemini API. Demis Hassabis followed up on X announcing that the upgrade set new records on “the most rigorous benchmarks in maths and science,” and India Today’s later coverage reported passing marks on Humanity’s Last Exam. If your team needs it now, the subscription tier is the only reliable path; API access requires applying and waiting.
Why Developers Should Care
Three takeaways. First, verifier design: natural-language verification plus permission to fail is a reusable agent architecture pattern for any high-stakes generation pipeline. Second, inference compute as a product variable: the 90% figure sits on a controllable quality-cost curve that belongs in your own eval and pricing design. Third, the differentiation battlefield is shifting from generic benchmarks to professional research workflows. Extending the acceleration and roadmap-divergence themes in our 2026 opening outlook, Deep Think bets on going deeper while the open-weight camp pushes prices down — and the split between those two routes will only get sharper.
Sources
- Accelerating Mathematical and Scientific Discovery with Gemini Deep Think — Google DeepMind
- Gemini 3 Deep Think: Advancing science, research and engineering — Google Blog
- Google Gemini 3 Deep Think AI scores passing marks in Humanity’s Last Exam — India Today
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
