10 mathematical problems, reversed in 24 hours.
This weekend, two AI giants, OpenAI and Anthropic, clashed head-on at the cutting edge of mathematics.
First, on August 1, OpenAI researcher Sébastien Bubeck announced that their unreleased next-generation flagship model, Astra, had proven 10 cutting-edge mathematical results in one go, also releasing 10 Lean certificates and 10 detailed solution explanations per problem.

Don't underestimate these 10 problems.
Their main conclusions had seen no progress for at least a decade, many had been dormant for decades, making them genuine open problems.

While not the deepest unsolved mysteries in mathematics, each one is a tough nut to crack.
Elliot Glazer, head of FrontierMath at Epoch AI, stated plainly: Just the paper on non-sophisticated groups alone would certainly make it into Annals of Mathematics, a top-tier journal recognized by the mathematical community.
The release of these 10 difficult problems sent shockwaves through the AI community, with some immediately declaring "the mathematical singularity has arrived."
Anthropic responded swiftly.
In less than 24 hours, researcher Levent Alpöge replied under Bubeck's post: "I completed half using Fable."

He then added that he had solved items 4, 5, 6, 7, and 8 from OpenAI's list, a total of 5 items.
The experimental conditions, Levent claimed, were clean: fully autonomous, generic prompts, no internet throughout, and a layer of protection was specifically added to prevent OpenAI's solutions from leaking into the context.

Of course, Levent hasn't yet released the complete proof details from Fable, stating he will upload them.
But if confirmed, the significance of this event changes:
The first-mover advantage OpenAI secured with its unreleased model Astra lasted only 24 hours. And the model catching up, Fable, is publicly available from Anthropic, accessible to everyone.
A "Mathematical First" Now Has a 24-Hour Shelf Life
User Chubby reposted: "Publicly available Fable replicated half of Astra's achievements."

A commenter joked: Looking at it this way, did OpenAI only secure a first announcement?
In the past, priority disputes over mathematical results typically played out over years.
Who thought of it first, who proved it first, who published first, separated by lengthy processes of submission, review, and revision.
Now, Astra hasn't even been formally released, yet the results it announced were half-matched by a public model in just 24 hours.
The value of being first remains, but its shelf life has drastically shortened.
Many netizens are already shouting "mathematical singularity," and the two AI giants are locked in a tight race.
Behind the $2000, There's an Even Pricier Bill
OpenAI also threw out a number: the cost of finding these 10 solutions was less than $2000.
According to researcher Noam Brown, this corresponds to the tokens required to search for these solutions, converted using the Sol API price.

In other words, this is just the marginal inference cost at the moment the 10 solutions were successfully produced.
Beyond that, there's the training and development of Astra itself, the problems attempted but not solved, the human effort to compile the arguments into a 249-page paper, the cost of formalizing each proof into Lean, and finally the cost of external mathematician review.
Noam himself admitted they had attempted other major problems without success, failing to crack any of the Millennium Prize Problems.
Nevertheless, $2000 still brings the marginal cost of an AI solving a cutting-edge problem down to rock bottom.
Some even joked whether OpenAI would announce 10 more mathematical breakthroughs tomorrow, or wait until next week and release 100 at once.

From "Benchmark Chasing" to "Mutual Verification"
Fable redoing the same batch of problems as Astra is more interesting than ordinary benchmark chasing.
Benchmarking tests known answers: the questions and standard answers are laid out, and you compare scores.
But two models, under conditions of no internet and leak protection, independently arriving at the same cutting-edge conclusion, tests something entirely different: how reliable is this result, can it be independently reproduced.
This is more like peer review.
Levent has currently only "made a statement" on X, not releasing the PDFs of those 5 proofs, prompts, complete run logs, token costs, or Lean certificates.
Netizen Haider questioned: Even if it's true, why use Fable to replicate Astra's work instead of directly tackling new open problems?

In fact, Fable isn't just riding Astra's coattails; it also solves new problems.
In July, Levent used Fable to produce a three-dimensional counterexample to the Jacobian Conjecture, and someone later wrote a 7-page algebraic verification material confirming that the explicit mapping did indeed refute the conjecture in three and higher dimensions.
For Achievements Most Can't Understand, Who Does the Accepting?
How important these 10 achievements are is very difficult for the vast majority to judge.
Ethan Mollick hit the nail on the head: For almost everyone on Earth, this is not just beyond ability, but beyond comprehension. We can only trust professional mathematicians to tell us how impressive it is.

With code, images, and chat, ordinary people can still personally experience where AI excels.
But the existence of non-sophisticated groups, counterexamples to Connes' Rigidity Conjecture, quantum parallel repetition theorems—these are understood by only a tiny minority of experts.
As AI capabilities advance towards the scientific frontier, what the public must rely on is not personal experience, but a long chain of trust. This state of "not even knowing how to be amazed" is happening simultaneously in more and more fields.
Mathematician Thomas Bloom cautioned that these achievements use mathematical theories accumulated over centuries and formal systems painstakingly built by mathematicians; simply describing it as "AI replacing mathematicians" is dishonest.
His criterion is: Does this proof teach us something new about the problem?
So the real question arises: When AI can produce cutting-edge mathematical proofs cheaply and in bulk, what should we use to measure its achievements?
The answer might not lie in its output end, but in the acceptance end: whether a proof can be read, verified, and have its significance judged ultimately determines its true value.
The scarcest ability is shifting from "producing proofs" to "understanding and accepting proofs."
When AI can prove ten difficult problems overnight, the real test is just beginning: As proofs are manufactured faster and faster, can our verification keep up?
References:
https://x.com/haider1/status/2083908869915611554
https://x.com/polynoamial/status/2083470822258467194
https://x.com/kimmonismus/status/2083950641978679363 https://x.com/Dr_Singularity/status/2083653287463571742
This article is from the WeChat public account "New Zhiyuan," author: ASI Apocalypse





