A one-page paper has tackled a tough nut that originally required a 44-page paper in a top journal to crack.
In 1991, a 44-page paper was published in the Annals of Mathematics.
The author was József Beck, a name well-known in the combinatorial mathematics community.
He is the proposer of the Beck-Fiala theorem, winner of the 1985 Fulkerson Prize, an invited speaker at the 1986 International Congress of Mathematicians, and one of the founders of discrepancy theory.
The paper was titled "Moduli of polynomials with zeros on the unit circle: a problem of Erdős," solving problem number 119 from Erdős's problem list (Erdős Problems).
Take a sequence of complex numbers on the unit circle as zeros, multiply them into a polynomial. Can the maximum modulus Mn of this polynomial on the unit circle be guaranteed to exceed n to the power of c for infinitely many n?
Beck proved it could. This problem was considered one of the toughest challenges of that era.

AI has claimed Erdős's bounty. Problem 119 originally asked three questions; the third, with a $100 reward, now has its status changed to SOLVED.
Thirty-five years later, this paper has been surpassed by a single page.
The one saying this is Thomas Bloom.
He is a mathematician, researcher at the University of Manchester, and the creator and maintainer of the website erdosproblems.com.
Over a thousand open problems left by Erdős are listed on this site, making it the primary destination for mathematicians worldwide tracking these problems.
For this reason, every time an AI company claims to have "solved an Erdős problem," it must first pass his scrutiny.
He wrote on X:
So far, GPT-5.6 Sol is the most interesting new model for mathematics I've seen. Every new proof submitted to erdosproblems.com that I've examined in detail has been correct and contained some interesting ideas.

Bloom also emphasized that this wasn't because the AI deployed more complex tools, but rather because we had previously assumed the problem was much harder than it actually is.
OpenAI President Greg Brockman reposted this tweet, saying:
Feels like a watershed moment for advancing mathematics. Scientific and medical breakthroughs that can truly improve human lives feel very close now.

Erdős's $100 Bounty, Claimed by AI
First, let's clarify the origin and details of this problem.
In 1957, Erdős posed this problem in a paper, asking three questions at once about the moduli of polynomials with zeros on the unit circle.
Over the next four decades, he revisited it in 1961, 1964, 1982, 1990, and 1997.
The first question was solved by Wagner in 1980.
The second question was solved by Beck in 1991, proving that there exists some c>0 such that max_{n≤N} M_n > N^c. This was that 44-page Annals of Mathematics paper.

Beck's paper was published in Annals of Mathematics, Volume 134, Issue 3, 1991.
The third question had a $100 bounty personally funded by Erdős, originating from a 1997 article of his.
Now, the bounty has been claimed by GPT-5.6 and a mathematician named Korsky. They solved the question Beck didn't touch using just one page.

Erdős proposed thousands of problems in his lifetime, offering personal bounties ranging from $25 to $10,000.
Bloom's assessment is: with Korsky's prompting, GPT produced a proof stronger than Beck's 44-page Annals paper using just one page of simple harmonic analysis techniques.
The crucial step, as Bloom later pointed out in the forum, was convolving a non-negative function and shifting it by 1 in time, which suddenly smoothed out the expression, making the previously hard part tractable.
He said this approach of putting operator norms into dual norms and inner products for estimation is one of the most common tricks in analysis.
Not a new tool, but an old tool applied where no one had thought to apply it.

Bloom clarified that Beck's Annals paper contains many important and interesting ideas, just not needed for this specific problem. He added: At least for me, this is quite surprising.
He also mentioned a detail: no one has calculated that constant; it's estimated to be very small. Beck himself didn't calculate it back then, and in his judgment, if the constant were worth mentioning, Beck would have calculated it.
So, this isn't AI overturning a top journal, but rather discovering a detour humans didn't need to take.
For mathematicians, this is more stimulating than "AI solved another problem."
Solving a problem just adds another tool. This time, the tool is correcting the tool user.
A Human Would Shrug and Walk Away; AI Tries the Next One
If this were an isolated case, it would merely be a neat anecdote.
GPT-5.6 Sol Ultra was fully released on July 9th. On July 10th, OpenAI announced that GPT-5.6 Sol Ultra had produced a complete proof of the Cycle Double Cover Conjecture.
This problem was independently proposed by Szekeres in 1973 and Seymour in 1979, remaining open for about fifty years.
The model was allocated 8 hours but used less than 1 hour, employing a method with 64 sub-agents working in parallel. The prompt and proof PDF are fully public.
Bloom was the first to provide a substantive evaluation.
He called it a very beautiful proof, then gave three adjectives:
Short. Elementary. Could have been discovered in the 1980s.
He also pointed out the underlying reason: he speculates there is a small, counterintuitive twist in the key step.
How would a human mathematician approach this problem?
They would first try the most natural approach, check the linear algebra, find it doesn't work, then shrug, think "didn't expect it to be that easy anyway," and walk away.
But AI doesn't get discouraged. It tries small variations until one works.
AI's strength is in not making emotional decisions to cut losses.
Looking back to April this year, GPT-5.4 Pro solved Erdős 1196 in 80 minutes using the von Mangoldt function.
Jared Lichtman from Oxford spent seven years on that problem. He later said that since 1935, everyone working on primitive set problems had become accustomed to the same opening move, which obscured a technical possibility that had been lying in plain sight for a full 90 years.
Tao's evaluation after reading it was even harsher: for decades, humans collectively took the wrong first step.
Human intuition is, of course, not a flaw. On the contrary, it's an efficiency tool: it saves us from a sea of doomed attempts. Without it, no one could reach the frontiers of a discipline in a finite lifetime.
The cost is that occasionally, it also saves the one attempt it shouldn't.
The Person Calling GPT-5.6 Interesting Said Something Else Last Year
Last October, Bloom publicly called out OpenAI.
At the time, OpenAI VP Kevin Weil reposted a post by Mark Sellke, writing: GPT-5 just found solutions to 10 previously unsolved Erdős problems, with progress on another 11, all open for decades.

Bloom responded with a sentence to characterize it: Dramatically misleading.
He said that these problems being marked as 'open' on the website only meant he personally wasn't aware of any paper that solved them. What GPT-5 did was dredge up literature he didn't know about.
And finding literature is not the same as constructing a proof.

The following scene is remembered by many.
LeCun chimed in nearby, subtly mocking the fact that his own planted mine exploded. Hassabis was more direct: This is so embarrassing.

Weil deleted the post.
OpenAI researcher Sébastien Bubeck eventually admitted they had only found solutions already existing in the literature.
When OpenAI announced in May this year that they had overturned a 1946 Erdős geometry conjecture, the announcement simultaneously included evaluations from Noga Alon, Melanie Wood, and Thomas Bloom.
Has AI Hit a Wall in Mathematics?
The debate under Bloom's post was also fascinating.
The pessimist faction was represented by user qrdl: "5.6's chain-of-thought output is pitifully little, no progress on problem 677, the CritPT evaluation on the physics side has hardly moved compared to 5.4."
His conclusion: AI has hit a wall in mathematics.

User qrdl posted on the forum, saying CritPT has barely moved since 5.4, and unless they found a way to game the benchmark, results are just about like this.
qrdl's argument wasn't emotional.
He believes large models are essentially static-weight token generators. When doing mathematics, they don't grow new neural connections in real-time like humans. True creative leaps require a flexible brain.
qrdl's feeling is: every time he starts a new session to attack 677, the model repeats the same old failed moves.
The counter-punch came quickly.
Nat Sothanaphan pointed out that CritPT is a physics test, not a mathematics benchmark. Using its curve alone to prove AI has hit a wall in math is cherry-picking favorable data.
His counter-evidence is: 5.6 shows improvements over 5.5 on a large set of evaluations including FrontierMath.
Another faction is the practical faction, represented by a user named old-bielefelder.

old-bielefelder reported that GPT-5.6 Sol, thinking for 14 minutes and reviewing for over 9 minutes, improved Pintz's 2018 exponent of 0.72 to 0.7195.
Not earth-shattering, but it did move, described by old-bielefelder himself as "the first taste of milk."
As a comparison, the night before, GPT-5.5 spent over an hour on the same problem without budging 0.72.
He then directly notified Pintz of the result.
The third voice also came from Nat Sothanaphan, who borrowed a saying from Bloom: Modern mathematics is a vast cathedral.
Reaching the very forefront of this cathedral itself consumes immense energy, and large models can already do this with superhuman efficiency; recent breakthroughs largely stem from this. Pushing further requires fluid intelligence.
Some say large models lack fluid intelligence; he believes this is wrong.
The steady rise in ARC series evaluations over recent years is evidence, with GPT-5.6 Sol achieving 7.8% at the highest reasoning tier. When this benchmark launched in March this year, the best score was 0.37%, while humans consistently score above 90%.

Official verification scores from ARC Prize. GPT-5.6 Sol's performance on ARC-AGI-3 rises steeply with reasoning tiers, from 0.3% at Low tier to 7.8% at Max tier.
His explanation is: the cathedral's foundation is too vast, vast enough to completely overshadow the small but rapidly growing fluid intelligence. So if you just look at the weight of the output, it seems like nothing is happening.
But the day fluid intelligence starts outweighing the foundation, it will suddenly become visible to the naked eye.
After two weeks of debate, what this argument truly brought out was not "Can AI do math?" but how to define the word "difficult."
Among the entries marked "unsolved" in mathematical history, some record the inherent difficulty of the problem, while others merely record the boundaries of human patience.
In the past, these two were mixed together, indistinguishable. Now, they are starting to be separable.
A problem hanging for decades without being claimed might not be because it's incredibly hard, but because no one was willing to try the thirtieth time on that path.
How many more problems are stuck simply because of human patience?
References:
https://annals.math.princeton.edu/1991/134-3/p03
https://www.erdosproblems.com/forum/thread/AI%20Contributions
https://arcprize.org/results/openai-gpt-5-6-sol
This article is from the WeChat public account "新智元" (New Zhiyuan), author: ASI启示录; editor: 元宇






