Google's Deep Think Dominates Eight-Language Olympiads, Autonomously Solves Four Unsolved Problems, Research Barriers Collapse

marsbitPublié le 2026-04-08Dernière mise à jour le 2026-04-08

Résumé

Google DeepMind's "Deep Think" AI system has demonstrated exceptional performance across eight languages in regional academic competitions, including mathematics and informatics Olympiads. It achieved perfect scores in Japanese and French contests, and high results in Chinese, Korean, Hindi, Vietnamese, Russian, and Portuguese exams. This multi-language capability aims to reduce linguistic barriers in scientific research, enabling non-English-speaking researchers to access advanced AI tools equally. Beyond competitions, Deep Think has solved four previously unsolved mathematical problems and contributed to breakthroughs in computer science, physics, and economics. It powers the Aletheia agent, which autonomously generates and verifies research-level mathematical solutions. Despite these achievements, the results are based on internal evaluations without third-party verification or detailed methodology disclosure. Google positions Deep Think as a "human intelligence multiplier," expanding AI's role in global scientific collaboration beyond English-dominated benchmarks.

"Deep Think has defeated/matched competitors in all competitions"!

Just now, Google DeepMind senior researcher Conglong Li posted 12 messages on the X platform, revealing an unprecedented scorecard.

One AI, the same brain, eight exam papers in different languages, all submitted with high scores.

Such results are rare for any model.

From IMO Gold Medals to Full Coverage of Regional Competitions

Deep Think's high scores across multiple leaderboards are not a sudden breakthrough but part of a nearly year-long evolution of capabilities.

First, it topped the most rigorous reasoning competitions.

In July 2025, Gemini Deep Think achieved the gold medal standard at the International Mathematical Olympiad (IMO) for the first time, scoring 35 out of 42 points. It also achieved similarly high-level performance at the ICPC World Finals around the same time.

These two achievements have been officially announced in the DeepMind blog.

Google DeepMind subsequently included these two results in its official blog, marking Deep Think's crossing of the "world-class competition threshold" in mathematics and programming.

Next, Deep Think began moving from "world-champion-level individual breakthroughs" to "systematic validation across languages, disciplines, and scenarios."

In February 2026, Google published three blog posts.

One introduced the Gemini 3.1 Pro model itself, one detailed a major upgrade to the Deep Think specialized reasoning mode, and one from the DeepMind scientific discovery team directly positioned Deep Think as a "human intelligence multiplier."

The upgraded Deep Think delivered a series of hard metrics:

48.4% on Humanity's Last Exam (without tool assistance), 84.6% on ARC-AGI-2 (officially verified by the ARC Prize Foundation), a Codeforces competitive programming Elo rating of 3455, and gold medal-level performance on the written portions of the 2025 International Physics and Chemistry Olympiads.

The strategy is very clear: first use world-class competitions like the IMO and ICPC to prove its powerful reasoning abilities, then use multi-language, regional competition, and cross-disciplinary Olympiad results to prove its general, deep reasoning ability that stably transfers across languages and domains.

Gemini Deep Think's capability evolution from IMO gold medals to PhD-level research acceleration

A Detailed Look at the 8-Language Scorecard

Now, let's take a closer look at this scorecard.

Japanese results are the most impressive.

2025 35th Japanese Mathematical Olympiad Finals (JMO Finals), perfect score.

ICPC Asia Japan Preliminary Contest, perfect score.

Among these, the JMO Finals score even exceeded the level corresponding to the top 80% of scores that year, meeting the official "gold medal equivalent" standard.

French results were also a perfect 100%.

The Chinese results are interesting.

At the 41st Chinese Mathematical Olympiad (CMO), Deep Think scored 86.3%, which is quite outstanding. But at the Chinese National Olympiad in Informatics (NOI), it only scored 63.3%.

The gap between 86.3% and 63.3% outlines the real boundaries of AI reasoning ability.

In math competitions, the model faces abstract deduction, proof construction, and multi-step reasoning, which happens to be Deep Think's strongest suit.

But in informatics competitions, the problem is not just "figuring it out," but also translating logic into executable code, controlling boundary conditions, considering complexity constraints, and avoiding implementation errors.

The former is closer to pure reasoning, while the latter requires "reasoning + algorithm design + engineering implementation" to be successful simultaneously.

In the other languages—Korean, Hindi, Vietnamese, Russian, Portuguese—Deep Think also achieved results that either defeated competitors or at least matched them.

Looking at Japanese, French, and Chinese together, the most unusual aspect this time is not necessarily scoring a perfect mark in any single subject, but rather that the same model, the same Deep Think reasoning system, delivered first-tier results on exam papers in multiple languages.

Is This Scorecard Reliable?

But there is a key omission:

Conglong Li did not list specific comparative data from competitors: all results come from Google evaluations. There is no independent third-party replication, no official certification from the competitions, and the evaluation methodology is completely undisclosed.

Was each problem attempted once or many times with the best score taken? How much computational power was used during reasoning? Was there any manual prompt engineering involved?

These details, which directly affect the credibility of the results, were also not mentioned.

Another easily overlooked point: these exams are all regional selection competitions, not international finals.

There is an order of magnitude difference in difficulty between regional competition problems and international finals.

The researcher explicitly stated that these results "will be included in the model card." As of publication, the model card has not been officially updated.

So, for now, this still seems like a scorecard graded by the examinee themselves, announced by themselves, and not yet stamped by the academic affairs office.

Multilingual Research Equity: The Overlooked Real Battlefield

Why did Google specifically invest effort in evaluating 8 different regional languages?

Current evaluations of AI reasoning ability are almost entirely based on English.

MATH, GSM8K, HumanEval, ARC-AGI... these are all in English.

Mathematicians, physicists, and engineers worldwide whose native language is not English must first overcome a language barrier when using AI research tools.

Google's selection of these 8 languages is not random.

Japanese, Korean, and Chinese cover East Asian research powerhouses; Hindi and Vietnamese cover emerging markets; French, Russian, and Portuguese cover Europe and South America.

Together, this represents the majority of global research output.

In its official blog, DeepMind positioned Deep Think as a "human intelligence multiplier," saying it can "handle knowledge retrieval and rigorous verification, allowing scientists to focus on conceptual depth and creative direction."

Combined with these multi-language results, the subtext of this statement is not hard to understand: this multiplier is not just for scientists who use English.

More notably is how far Deep Think has already gone in research落地 (landing/application).

DeepMind announced a mathematical research agent called Aletheia, powered by Deep Think, capable of autonomously generating, verifying, and revising solutions to research-level mathematical problems.

Aletheia, driven by Deep Think, capable of iterative generation, verification, and correction for research-level mathematical problems

Aletheia has already contributed to multiple research papers, one of which was completed entirely autonomously by the AI, calculating specific structural constants in arithmetic geometry.

Furthermore, in a semi-autonomous evaluation of 700 open mathematical problems, it independently solved 4 previously unsolved problems.

The Gemini Deep Think mode also shows great potential in computer science, physics, economics, and other fields.

In computer science, Deep Think helped refute a conjecture that had remained open for a decade; in physics, it found a new analytical solution for gravitational radiation from cosmic strings; in economics, it extended an auction theory theorem.

Schematic diagram of the AI reasoning process, showing how large-scale exploration of the solution space at the network layer is aggregated into structured reasoning and confirmed through automated and manual verification.

By collaborating with experts to solve 18 research challenges, the advanced version of Gemini Deep Think helped break through long-standing bottlenecks in algorithms, machine learning and combinatorial optimization, information theory, and economics.

This goes far beyond "solving competition problems."

While competitors are still competing on English benchmark leaderboards, Google has already found a new battlefield in the "AI research accelerator" field.

The most important thing about this is not the scores; the real signal behind it is: the language barrier for AI research tools is being treated as an engineering problem to be solved.

If this path succeeds, scientists conducting research in Japanese, Korean, Chinese, Hindi, and other languages will, for the first time, stand on the same starting line as native English speakers.

This time, Google has laid its cards on the table.

As for which competitors will follow suit, we believe we will see soon.

References:

https://blog.google/intl/ja-jp/company-news/technology/gemini-31-pro-gemini-31-pro-deep-think/%20

https://deepmind.google/blog/accelerating-mathematical-and-scientific-discovery-with-gemini-deep-think/%20

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/%20

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-deep-think/

This article is from the WeChat public account "新智元" (New Zhiyuan), author: 新智元

Questions liées

QWhat is the key achievement of Google's Deep Think AI model as reported in the article?

ADeep Think achieved top-tier results in eight different language versions of academic competitions, including perfect scores in Japanese and French math and programming contests, and high performance in Chinese, Korean, Hindi, Vietnamese, Russian, and Portuguese exams.

QWhich specific world-class competitions did Deep Think first demonstrate its reasoning capabilities in?

ADeep Think first demonstrated its reasoning capabilities by reaching gold medal standards in the International Mathematical Olympiad (IMO) with a score of 35 out of 42 in July 2025, and achieving similarly high performance in the ICPC World Finals.

QWhat is the significance of Deep Think's performance across multiple languages according to the article?

AIts performance across multiple languages signifies a breakthrough in breaking down language barriers in AI research tools, potentially allowing non-English speaking scientists worldwide to access advanced AI research assistance on equal footing.

QWhat are some research breakthroughs mentioned that were achieved using Deep Think?

ADeep Think autonomously solved 4 previously unsolved mathematical problems, refuted a decade-old conjecture in computer science, found new analytical solutions for cosmic string gravitational radiation in physics, and extended an auction theory theorem in economics.

QWhat concerns does the article raise about the reliability of Deep Think's reported results?

AThe article notes that all results are from internal Google evaluations without third-party verification, official contest authentication, or disclosure of testing methods such as attempt counts, computational resources used, or potential human prompt engineering involvement.

Lectures associées

Interview sur l'ère de l'IA, la révolution industrielle et la civilisation future — Zhang Dingwen : L’avenir n’appartient pas à ceux qui se contentent de suivre

Dans cet entretien, l'entrepreneur Zhang Dingwen partage sa vision sur l'innovation, l'ère de l'IA et la construction d'entreprises durables. Il souligne que le succès ne réside pas dans la poursuite des tendances éphémères, mais dans la compréhension des mouvements profonds de l'époque et de l'évolution à long terme des besoins humains. Il retrace son parcours, de ses débuts dans l'internet étudiant à ses expériences entrepreneuriales, en tirant des leçons fondamentales : la valeur utilisateur ne se traduit pas automatiquement en valeur commerciale, et l'échec est une opportunité d'amélioration cognitive. Pour lui, l'entrepreneuriat consiste à affiner constamment sa perception du monde, à poser les bonnes questions et à comprendre les causes profondes derrière les succès et les échecs. Zhang Dingwen évolue d'une mentalité centrée sur le produit vers une pensée systémique et une quête de "points d'entrée" stratégiques. Il voit dans le matériel intelligent, comme les montres connectées, non pas de simples gadgets, mais de futures plates-formes cruciales reliant les utilisateurs aux services, aux données et à des écosystèmes entiers. Ces dispositifs combinent attributs technologiques, financiers, sociaux et même de mode, visant à établir une relation durable plutôt qu'une transaction unique. Enfin, il élargit la perspective au-delà de l'entreprise : les organisations véritablement grandes ne se contentent pas de concurrencer sur les produits ou les modèles économiques ; elles participent à façonner l'avenir, définissent de nouvelles règles et contribuent au progrès de la civilisation. Leur mission ultime est de résoudre les problèmes de leur temps, de bâtir une confiance inébranlable et de laisser un héritage de valeur qui transcende la richesse matérielle. L'avenir, conclut-il, appartient non pas à ceux qui courent le plus vite, mais à ceux qui gardent une capacité d'apprentissage et d'adaptation constante.

marsbitIl y a 3 mins

Interview sur l'ère de l'IA, la révolution industrielle et la civilisation future — Zhang Dingwen : L’avenir n’appartient pas à ceux qui se contentent de suivre

marsbitIl y a 3 mins

Base sous pression

Le fondateur de Base, Jesse Pollak, a récemment reconnu que la stratégie de la plateforme axée sur les jetons sociaux et pour créateurs était une erreur. Alors que la nouvelle Robinhood Chain gagne rapidement du terrain avec des volumes d'échange élevés, Base fait face à des critiques croissantes concernant son manque de décentralisation, au point que L2BEAT envisage de rétrograder son statut. De plus, un incident récent impliquant le fondateur de Coinbase, Brian Armstrong, et un jeton meme a suscité des moqueries de la communauté, accentuant la pression. Malgré cela, Base maintient la TVL la plus élevée parmi les L2 (près de 12 milliards de dollars) et un rôle clé dans les paiements automatisés. Cependant, l'émergence de Robinhood Chain, notamment avec son approche des actifs tokenisés, force Base à reconsidérer ses priorités. Pour rester compétitive face à cette concurrence accrue et à de possibles nouveaux entrants, Base doit urgemment résoudre ses problèmes techniques de longue date et regagner la confiance des utilisateurs, au-delà de la simple poursuite d'une croissance à court terme.

Foresight NewsIl y a 7 mins

Foresight NewsIl y a 7 mins

La Maison Blanche fait une concession pour lever les obstacles éthiques, la "Clarity Act" va-t-elle saisir la dernière fenêtre avant la suspension des travaux ?

Le gouvernement Trump aurait accepté d'intégrer des dispositions éthiques au "Clarity Act", un projet de loi phare sur la structure du marché des actifs numériques, ce qui pourrait lever le dernier obstacle majeur à son adoption. Le texte mis à jour a été soumis à des sénateurs républicains. Parallèlement, Patrick Witt, conseiller clé de la Maison Blanche, restera en poste pour superviser la dernière phase des négociations. Le "Digital Asset Market Clarity Act" vise à établir un cadre réglementaire fédéral unifié pour clarifier la classification des actifs numériques (marchandises numériques, actifs de contrat d'investissement, stablecoins) et à répartir les compétences entre la SEC et la CFTC, mettant fin à des années de flou juridique. Après des compromis sur les revenus des stablecoins et la régulation de la DeFi, la principale divergence restante concernait les conflits d'intérêts potentiels des responsables gouvernementaux. La concession de la Maison Blanche sur cette clause éthique ouvre une voie possible pour un vote au Sénat avant la période de récess estivale du Congrès, qui débute mi-août. Les acteurs de l'industrie voient cela comme un tournant potentiel, offrant enfin une clarté réglementaire susceptible d'attirer les institutions traditionnelles et de consolider le marché américain des actifs numériques.

Odaily星球日报Il y a 11 mins

La Maison Blanche fait une concession pour lever les obstacles éthiques, la "Clarity Act" va-t-elle saisir la dernière fenêtre avant la suspension des travaux ?

Odaily星球日报Il y a 11 mins

Le hack de 515 millions de NIGHT sur Midnight envoie le token en chute de 32 % – Le support à 0,015 $ tiendra-t-il ?

L'année 2026 a été marquée par une recrudescence des piratages dans le secteur cryptographique, avec plus d'un milliard de dollars volés à ce jour. Le réseau Midnight est la dernière victime en date, subissant une exploitation sur le pont inter-chaînes Wanchain entre Cardano et BNB. Un contrat datant de deux ans, contenant 515 millions de jetons NIGHT, a été drainé. Les fonds ont ensuite été vendus sur des DEX Cardano. Suite à cet incident, le jeton NIGHT s'est effondré de 32%, atteignant un plus bas historique à 0,015 dollar avant un léger rebond à 0,019 dollar. La capitalisation boursière a chuté de 27% tandis que le volume d'échanges a explosé de 829%, signe d'une intense pression vendeuse amplifiée par les ventes de l'attaquant. L'Indice de Force Relative (RSI) est tombé à 17, confirmant un territoire de survente. La Fondation Midnight a précisé que le réseau principal n'était pas compromis, l'incident étant isolé aux opérations du pont. Dans l'immédiat, le sentiment reste fortement baissier. Si cette tendance persiste, le NIGHT pourrait évoluer sous la barre des 0,02 dollar, avec le niveau de 0,015 dollar comme prochain support.

ambcryptoIl y a 17 mins

Le hack de 515 millions de NIGHT sur Midnight envoie le token en chute de 32 % – Le support à 0,015 $ tiendra-t-il ?

ambcryptoIl y a 17 mins

Une Transparence des Réserves Construite Grâce à une Exploitation Continue : Matrixdock Fête Deux Ans de Vérification Indépendante

Matrixdock a achevé son quatrième audit semestriel indépendant consécutif des réserves avec Bureau Veritas, marquant ainsi deux ans de vérification continue. Pour la première fois, l'audit s'est étendu au produit d'argent tokenisé (XAGm) en plus de l'or (XAUm). L'audit physique a confirmé que 574 barres d'or et d'argent, détenues dans des installations de Malca-Amit et Brink's, correspondent exactement aux enregistrements. Les réserves auditées se montent à environ 66,09 millions de dollars pour l'or et 4,04 millions de dollars pour l'argent, alignées sur l'offre de tokens en circulation. Cette transparence récurrente, renforcée par des états mensuels et des outils de vérification sur la chaîne, constitue la base de confiance permettant à ces actifs de réserve d'être utilisés comme collatéral et dans les infrastructures financières décentralisées.

TheNewsCryptoIl y a 50 mins

Une Transparence des Réserves Construite Grâce à une Exploitation Continue : Matrixdock Fête Deux Ans de Vérification Indépendante

TheNewsCryptoIl y a 50 mins

Trading

Spot

Google's Deep Think Dominates Eight-Language Olympiads, Autonomously Solves Four Unsolved Problems, Research Barriers Collapse

Résumé

From IMO Gold Medals to Full Coverage of Regional Competitions

A Detailed Look at the 8-Language Scorecard

Is This Scorecard Reliable?

Multilingual Research Equity: The Overlooked Real Battlefield

Questions liées

Lectures associées

Interview sur l'ère de l'IA, la révolution industrielle et la civilisation future — Zhang Dingwen : L’avenir n’appartient pas à ceux qui se contentent de suivre

Base sous pression

La Maison Blanche fait une concession pour lever les obstacles éthiques, la "Clarity Act" va-t-elle saisir la dernière fenêtre avant la suspension des travaux ?

Le hack de 515 millions de NIGHT sur Midnight envoie le token en chute de 32 % – Le support à 0,015 $ tiendra-t-il ?

Une Transparence des Réserves Construite Grâce à une Exploitation Continue : Matrixdock Fête Deux Ans de Vérification Indépendante

Trading

Catégories populaires

Tags tendances