Terence Tao's Remote Dialogue with Wang Hong: Mathematics Must Learn to 'Digest' AI

marsbitPublicado em 2026-08-20Última atualização em 2026-08-20

Resumo

Two Fields Medalists, Terence Tao and Maryna Viazovska, share a common view on the flood of AI-generated proofs in mathematics: the community must learn to "digest" them. While AI can rapidly produce and even formally verify proofs, Tao argues that a correct proof is only the first step. For a result to become usable knowledge, it must be understood, absorbed, and integrated into the existing mathematical framework by human mathematicians. He illustrates this by spending several days "digesting" an AI-assisted proof of the long-standing Sendov conjecture. His process involved tracing sources, identifying key ideas, simplifying the argument, and rewriting it into a human-readable narrative. This digestion not only verified the proof but also revealed it could solve a stronger conjecture and be presented with more elementary tools. Tao criticizes the current race for priority based solely on who announces a proof first, often skipping verification and explanation. He proposes shifting value to the crucial work of interpreting, reviewing, and consolidating results. In his recent ICM talk, he emphasized that if authors cannot explain their AI-generated proof to peers, it shouldn't be published. To facilitate this new workflow, Tao introduced "Palomar," a registry for Lean-verified results. It serves as a digestion hub, cataloging problems, proof code, and AI involvement, aiming to coordinate efforts and ensure completeness. Under this model, full credit for a result would requ...

Too many AI-generated proofs isn't necessarily a good thing... the mathematics community can hardly keep up with reading them.

But there's good news—

In the face of this crisis, two generations of Fields Medalists, Terence Tao and Wang Hong, have offered nearly identical judgments:

Digest.

The School of Mathematical Sciences at Peking University recently released a lengthy 10,000-word interview with Wang Hong, where she offered quite pertinent advice:

The mathematics community cannot ignore the counterexamples or proofs generated by AI; we need to learn to use them, but also understand and digest them.

Terence Tao, on the other hand, spent several days thoroughly digesting a proof that involved AI from start to finish.

This proof solves the Sęndov conjecture, which had been pending for 67 years.

This sounds somewhat counterintuitive. If AI is here to help mathematicians solve problems and the answer is already out, why digest it again?

The answer from 'whistleblower' Terence Tao is:

A correct proof is only the first hurdle for a mathematical result; it must be understood and absorbed by professional mathematicians to be fully usable.

To use an analogy, traditional authors are like responsible parents, fully accompanying a proof from birth to publication. But moving forward, working with AI, this relationship may be more like 'adoption':

Someone uses AI to advance a proof to the generation or verification stage, then formally hands it off to another group for professional interpretation, organization, and publication.

Mathematicians still have much to do.

Digestion is the Next Step in Solving Mathematics' Century-Old Crisis

To put it bluntly, this is also Terence Tao's response to the question he left in his ICM lecture a month ago:

If answers can be mass-produced by AI, what is mathematical research really pursuing?

OpenAI's new model Astra solved 10 challenging problems in mathematics and theoretical computer science in a row, spanning high-dimensional geometry, coding theory, group theory, and more.

A neurosurgeon without formal higher mathematics training had ChatGPT 5.6 autonomously run for about 16 hours, producing a key proof for Crouzeix's conjecture in numerical linear algebra.

......

The high wall of mathematics seems to have had a back door chiseled open by AI.

Even laypeople can use it to break through cutting-edge problems. Is mathematics about to get its "next civil engineering" script? (doge)

Not at all! In fact, a complete mathematical achievement is never just about 'obtaining a proof'.

Terence Tao breaks it down into five stages:

First, generate the argument; then verify its correctness; next, explain the idea to peers, letting it undergo publication and community verification; finally, reorganize the results, find their relationship to old theories, and solidify them into standard knowledge for future use.

AI currently excels at accelerating the first two steps. The further along you go, the more you need mathematicians to judge what's truly core and what the results mean.

And the Sęndov conjecture Tao just processed is almost a 1:1 live demonstration of this workflow.

The conjecture studies the distance between polynomial zeros and critical points. Lower-degree cases and sufficiently high-degree cases had already been solved, but the middle range remained vacant for a long time.

A few days ago, math enthusiast Lech Mazur used AI to fill this gap, providing a formalized proof verified by Lean.

It seemed the story should end there. But Terence Tao noted that this original proof hadn't been organized into mathematical text suitable for human reading and publication, meaning the mathematical result was not yet usable.

So he spent several days, with the assistance of both ChatGPT and pen-and-paper derivation, conducting a complete digestion of the result:

Tracing literature sources → Extracting the identities that truly mattered → Removing detours → Rewriting the machine-discovered argument into a version showing a clear main line.

Reading it this way, he actually uncovered something new.

The reorganized argument not only solves the Sęndov conjecture but can also cover the stronger Phelps–Rodriguez conjecture. The core tools required are also much more elementary than what the original formalization presented, compressing the Lean code from about 90,000 lines to 15,000 lines.

In other words, digestion is not just a re-verification of an AI proof, but also an effective way to broaden mathematical achievements.

More importantly, Tao wants to use digestion to overturn the old rule in mathematics of "whoever announces first gets priority."

The attribution of mathematical results was first based on journal publication date, then changed to arXiv timestamp.

Now AI proofs come too fast, and even arXiv is considered too slow by some. Some results are thrown onto social platforms immediately after generation. In the race to be first, verification and interpretation are all skipped, ultimately creating an artificial burden for the mathematics community.

Therefore, in his ICM speech, Tao argued that the mathematics community should stop blindly celebrating "the first person to give a proof" and elevate the status of "digesting and organizing proofs."

Work like explaining proofs, peer review, and organizing results into classical theory should also be valued; not only producing new theorems counts as merit.

If the paper's authors themselves cannot give a competent lecture at a professional level to explain the results clearly, no matter how correct the AI verification is, this achievement should not be published.

If even humans can't understand the proof, how can it be considered a complete result?

So Tao calls on the academic community to calmly sit down, engage in a profound discussion about AI's capability limits and the values of mathematics, while emphasizing those parts of mathematical work that cannot be replaced by machines.

One More Thing

Meanwhile, Terence Tao has also publicly released a real registry for Lean verification results — Palomar.

Palomar clearly registers the problem statement, proof code, AI involvement method, and version information. After independent kernel re-evaluation, it gathers AI proofs scattered across GitHub, social media, and news in one place.

It will serve as a digestion relay station between verification and publication, letting future researchers know where AI and peers have reached and what work remains to be done.

Teams interested in the same problem can collaborate via Palomar.

Notably, Palomar does not judge the importance of a particular result, nor does it replace peer review.

The attribution of specific results will be determined in the future by the chronological order of four key stages:

Generative achievement documenting how the proof was produced, e.g., AI chat logs;

Verification achievement confirming logical correctness, e.g., Lean code;

Interpretation achievement clearly explaining the idea, e.g., public lecture;

Publication achievement formally presenting the result, e.g., paper.

In other words, priority is granted to whoever first submits a complete set of all achievements.

Currently, Terence Tao's digested version of the Sęndov conjecture has become one of the first entries in Palomar's archives.

References:

[1]https://arxiv.org/abs/2608.16753

[2]https://mathstodon.xyz/@tao/117123957813926355

[3]https://www.youtube.com/watch?v=M0--ZH1lOzg

[4]https://mp.weixin.qq.com/s/w-UfVMmwTZqXjNsC81CzYw

This article is from the WeChat public account "QbitAI", author: Focus on Frontier Technology

Perguntas relacionadas

QAccording to the article, what is the main challenge faced by the mathematics community regarding AI-generated proofs?

AThe mathematics community is struggling to keep up with the sheer volume of AI-generated proofs and counterexamples, creating a need to 'digest' this information effectively.

QHow did Terence Tao and June Huh suggest the mathematics community should approach AI-generated mathematical content?

ABoth Terence Tao and June Huh emphasized the need for the mathematics community to 'digest' AI outputs—meaning they must learn to use, understand, and integrate these AI-generated proofs or counterexamples into the field's knowledge base.

QWhat are the five stages of a complete mathematical result as outlined by Terence Tao in the article?

ATerence Tao outlined five stages: 1) Generating the argument, 2) Verifying its correctness, 3) Explaining the idea to colleagues, 4) Getting it published and passing community scrutiny, and 5) Reorganizing the result, relating it to existing theory, and solidifying it as standard, reusable knowledge.

QWhat is 'Palomar' and what is its intended purpose, as mentioned in the article?

A'Palomar' is a registry, proposed by Terence Tao, for formalized (e.g., Lean-verified) mathematical results. It serves as a 'digestion relay station' between verification and publication, providing a centralized record of AI-assisted proofs, their source code, and verification details to prevent duplication and facilitate collaboration.

QAccording to Tao's proposed new framework, what determines the attribution of priority for a mathematical result in the age of AI?

APriority would be determined by who first completes a full set of four milestones: the generative work (e.g., AI chat logs), the verification work (e.g., Lean code), the expository work (e.g., a clear public lecture), and the published work (e.g., a formal paper). The complete suite, not just the initial proof, grants priority.

Leituras Relacionadas

Gold Diggers in Prediction Markets: From Competing for Trading Entrances to Competing for Outcome Definition Rights

The report identifies a shift in prediction market competition from front-end user acquisition to back-end infrastructure, specifically the "outcome layer." This layer encompasses the standardized services for rule comparison, evidence verification, outcome confirmation, and payment triggering. Analysis shows that while a tiny fraction (0.487%) of markets face disputes, they account for a significant share (8.64%) of traded volume. This highlights the financial impact of rule uncertainty, which creates trading alpha but limits strategy capacity due to shallow order books. The larger opportunity lies in productizing these backend functions. Services like automated settlement (e.g., HIP-4), AI-assisted evidence processing, and external data oracles (e.g., Pyth, Chainlink) are becoming reusable, cross-platform infrastructure. This is creating a "second profit pool" separate from trading fees. Current observable revenue for this outcome layer is estimated at $15-37 million annually. If applied to the entire existing market, this could expand to $64-161 million. In a mature state, modeled after existing commercial models like Azuro's, annual revenue potential could reach approximately $456 million. While the industry logic is forming, pure-play investment assets are still early. Platform equities (e.g., Kalshi, Polymarket) price in broad growth, not just the outcome layer. Tokens like HYPE have minimal fee contribution from related products, and ICE's exposure is too small relative to its total business. The key is to track early projects that achieve cross-platform adoption and convert usage into attributable, recurring revenue. The most significant alpha may emerge before the ideal investment target is fully established.

marsbitHá 29m

Gold Diggers in Prediction Markets: From Competing for Trading Entrances to Competing for Outcome Definition Rights

marsbitHá 29m

Bitcoin Surpasses $72k. What Happens Next?

Bitcoin exceeded $72,000, reaching its highest level since early summer after a $8,000 surge. The broader crypto market followed, with total market capitalization rising 10%. Analysts attribute the rally to a combination of factors: the U.S. Treasury's decision to increase purchases of long-term government bonds, which lowers yields and makes riskier assets like crypto more attractive, and positive remarks about cryptocurrency from former President Donald Trump. Experts note that while Trump's comments provided an emotional boost, the Treasury's action was a more fundamental driver. However, they caution that significant sustained growth is not assured. High U.S. inflation limits the Federal Reserve's ability to ease policy, and network activity within crypto ecosystems remains low. Competition for capital from sectors like semiconductors and AI also persists. Analysts present mixed short-term forecasts. Some see resistance around $72,000; a breakout could target $76,000-$80,000, but a moderate, localized rise within a longer-term downward trend is deemed more likely. Others identify $69,000-$70,000 as a key support zone. A sustained move above $72,000 could open a path toward $75,000-$78,000, or even $80,000-$82,000 under a strong bullish scenario, with support expected around $66,000-$67,000 on any pullback. The overall outlook is cautiously optimistic, with the recent surge signaling improved sentiment, but a confirmed upward trend requires prices to solidify above current levels.

cryptonews.ruHá 32m

Bitcoin Surpasses $72k. What Happens Next?

cryptonews.ruHá 32m

Optimism Developers' Vote Deprives Users of $49 Million Airdrop

On August 19th, the Optimism community approved a proposal to reallocate 546.9 million OP tokens (worth ~$49.7 million) from a reserve for future user airdrops into a new Strategic Ecosystem Fund for development. The vote passed with 17.97 million OP for and 10.93 million against, with a decisive last-minute vote from developer group Test in Prod securing the majority. The funds will now be managed by the Optimism Foundation to foster ecosystem growth, targeting partnerships with blockchains, protocols, and financial firms, and attracting major brands. This marks a strategic shift from broad user acquisition via airdrops to focusing on corporate clients and enterprise solutions (OP Enterprise). The decision does not affect past airdrops, which distributed 269.1 million OP over five rounds. The proposal faced criticism. L2BEAT and delegate cp0x argued the Foundation gains overly broad discretion without sufficient detail on fund use, results measurement, or benefit to OP holders. Critics also noted these tokens were publicly reserved for future user distributions, and their reallocation removes the need for separate community votes on individual Foundation deals. Test in Prod defended its support, stating the reserve is needed for competitive corporate deals where premature disclosure could be detrimental, while urging the Foundation to later disclose aggregate spending and outcomes.

cryptonews.ruHá 33m

Optimism Developers' Vote Deprives Users of $49 Million Airdrop

cryptonews.ruHá 33m

Trading

Spot
活动图片