OpenAI Solves 10 Mathematical Problems, Fable 'Replicates' 5 in 24 Hours

marsbitPublished on 2026-08-05Last updated on 2026-08-05

Abstract

OpenAI and Anthropic engaged in a rapid, high-stakes competition at the cutting edge of mathematics this weekend. On August 1, OpenAI researcher Sébastien Bubeck announced that their next-generation model Astra had autonomously solved 10 longstanding, open mathematical problems, providing Lean proofs and solution breakdowns. The problems, untouched for years, are considered significant; one result on non-sofic groups is deemed worthy of a top mathematics journal. The estimated marginal cost for these solutions was under $2,000. Within 24 hours, Anthropic researcher Levent Alpöge responded, stating he had independently used the publicly available model Fable to solve 5 of the 10 problems (#4-8), under clean conditions without internet access and with safeguards against data leakage. This dramatically shortens the "shelf life" of a mathematical discovery, shifting priority from years to potentially a day. The event is seen less as simple benchmarking and more as a form of peer review, testing the reliability and independent reproducibility of AI-generated proofs. The episode raises critical questions about validation in the age of AI. As these models can now produce complex proofs at low marginal cost, their outputs are often beyond public comprehension. The true challenge shifts from generating proofs to verifying, understanding, and judging their significance—a task that remains a deeply human and expert-driven endeavor. The ability to critically evaluate AI's mathematica...

10 mathematical problems, reversed in 24 hours.

This weekend, two AI giants, OpenAI and Anthropic, clashed head-on at the cutting edge of mathematics.

First, on August 1, OpenAI researcher Sébastien Bubeck announced that their unreleased next-generation flagship model, Astra, had proven 10 cutting-edge mathematical results in one go, also releasing 10 Lean certificates and 10 detailed solution explanations per problem.

Don't underestimate these 10 problems.

Their main conclusions had seen no progress for at least a decade, many had been dormant for decades, making them genuine open problems.

While not the deepest unsolved mysteries in mathematics, each one is a tough nut to crack.

Elliot Glazer, head of FrontierMath at Epoch AI, stated plainly: Just the paper on non-sophisticated groups alone would certainly make it into Annals of Mathematics, a top-tier journal recognized by the mathematical community.

The release of these 10 difficult problems sent shockwaves through the AI community, with some immediately declaring "the mathematical singularity has arrived."

Anthropic responded swiftly.

In less than 24 hours, researcher Levent Alpöge replied under Bubeck's post: "I completed half using Fable."

He then added that he had solved items 4, 5, 6, 7, and 8 from OpenAI's list, a total of 5 items.

The experimental conditions, Levent claimed, were clean: fully autonomous, generic prompts, no internet throughout, and a layer of protection was specifically added to prevent OpenAI's solutions from leaking into the context.

Of course, Levent hasn't yet released the complete proof details from Fable, stating he will upload them.

But if confirmed, the significance of this event changes:

The first-mover advantage OpenAI secured with its unreleased model Astra lasted only 24 hours. And the model catching up, Fable, is publicly available from Anthropic, accessible to everyone.

A "Mathematical First" Now Has a 24-Hour Shelf Life

User Chubby reposted: "Publicly available Fable replicated half of Astra's achievements."

A commenter joked: Looking at it this way, did OpenAI only secure a first announcement?

In the past, priority disputes over mathematical results typically played out over years.

Who thought of it first, who proved it first, who published first, separated by lengthy processes of submission, review, and revision.

Now, Astra hasn't even been formally released, yet the results it announced were half-matched by a public model in just 24 hours.

The value of being first remains, but its shelf life has drastically shortened.

Many netizens are already shouting "mathematical singularity," and the two AI giants are locked in a tight race.

Behind the $2000, There's an Even Pricier Bill

OpenAI also threw out a number: the cost of finding these 10 solutions was less than $2000.

According to researcher Noam Brown, this corresponds to the tokens required to search for these solutions, converted using the Sol API price.

In other words, this is just the marginal inference cost at the moment the 10 solutions were successfully produced.

Beyond that, there's the training and development of Astra itself, the problems attempted but not solved, the human effort to compile the arguments into a 249-page paper, the cost of formalizing each proof into Lean, and finally the cost of external mathematician review.

Noam himself admitted they had attempted other major problems without success, failing to crack any of the Millennium Prize Problems.

Nevertheless, $2000 still brings the marginal cost of an AI solving a cutting-edge problem down to rock bottom.

Some even joked whether OpenAI would announce 10 more mathematical breakthroughs tomorrow, or wait until next week and release 100 at once.

From "Benchmark Chasing" to "Mutual Verification"

Fable redoing the same batch of problems as Astra is more interesting than ordinary benchmark chasing.

Benchmarking tests known answers: the questions and standard answers are laid out, and you compare scores.

But two models, under conditions of no internet and leak protection, independently arriving at the same cutting-edge conclusion, tests something entirely different: how reliable is this result, can it be independently reproduced.

This is more like peer review.

Levent has currently only "made a statement" on X, not releasing the PDFs of those 5 proofs, prompts, complete run logs, token costs, or Lean certificates.

Netizen Haider questioned: Even if it's true, why use Fable to replicate Astra's work instead of directly tackling new open problems?

In fact, Fable isn't just riding Astra's coattails; it also solves new problems.

In July, Levent used Fable to produce a three-dimensional counterexample to the Jacobian Conjecture, and someone later wrote a 7-page algebraic verification material confirming that the explicit mapping did indeed refute the conjecture in three and higher dimensions.

For Achievements Most Can't Understand, Who Does the Accepting?

How important these 10 achievements are is very difficult for the vast majority to judge.

Ethan Mollick hit the nail on the head: For almost everyone on Earth, this is not just beyond ability, but beyond comprehension. We can only trust professional mathematicians to tell us how impressive it is.

With code, images, and chat, ordinary people can still personally experience where AI excels.

But the existence of non-sophisticated groups, counterexamples to Connes' Rigidity Conjecture, quantum parallel repetition theorems—these are understood by only a tiny minority of experts.

As AI capabilities advance towards the scientific frontier, what the public must rely on is not personal experience, but a long chain of trust. This state of "not even knowing how to be amazed" is happening simultaneously in more and more fields.

Mathematician Thomas Bloom cautioned that these achievements use mathematical theories accumulated over centuries and formal systems painstakingly built by mathematicians; simply describing it as "AI replacing mathematicians" is dishonest.

His criterion is: Does this proof teach us something new about the problem?

So the real question arises: When AI can produce cutting-edge mathematical proofs cheaply and in bulk, what should we use to measure its achievements?

The answer might not lie in its output end, but in the acceptance end: whether a proof can be read, verified, and have its significance judged ultimately determines its true value.

The scarcest ability is shifting from "producing proofs" to "understanding and accepting proofs."

When AI can prove ten difficult problems overnight, the real test is just beginning: As proofs are manufactured faster and faster, can our verification keep up?

References:

https://x.com/haider1/status/2083908869915611554

https://x.com/polynoamial/status/2083470822258467194

https://x.com/kimmonismus/status/2083950641978679363 https://x.com/Dr_Singularity/status/2083653287463571742

This article is from the WeChat public account "New Zhiyuan," author: ASI Apocalypse

Related Questions

QWhat did OpenAI's unreleased model Astra reportedly achieve in the field of mathematics?

AOpenAI's unreleased model Astra reportedly proved 10 previously unsolved, open mathematical problems and provided Lean certificates and solution analyses for each.

QWhat significant action did Anthropic's model Fable take in response to OpenAI's announcement, and under what conditions?

AWithin 24 hours, Anthropic's publicly available model Fable reportedly proved 5 of the same 10 problems independently, using a general prompt and no internet access to prevent contamination from OpenAI's solutions.

QAccording to the article, what does the $2000 cost figure from OpenAI represent regarding the math proofs?

AThe $2000 cost figure represents the marginal inference cost of generating the 10 successful proofs using OpenAI's API, but does not include the costs of training the model, failed attempts, human labor for formatting, or external mathematical review.

QWhy does the article suggest that Fable's reproduction of Astra's work is more significant than standard benchmark competition?

ABecause it involves independently arriving at the same novel, unsolved conclusions under controlled conditions, which serves as a form of peer review to verify the reliability and reproducibility of the results, rather than just testing known answers.

QWhat core challenge does the article identify regarding the validation of AI-generated mathematical proofs?

AThe core challenge is that the ability to understand, verify, and assess the significance of highly complex, frontier mathematical proofs is extremely scarce. As AI produces proofs faster, the ability to validate them—not just generate them—becomes the true bottleneck and measure of value.

Related Reads

DeFi Sector Bounces Back Strongest: Which High-Revenue Projects Offer Entry Opportunities?

DeFi Sector Leads Recovery: Which High-Revenue Projects Are Worth Watching? DeFi has been one of the most active sectors during the recent market rebound. Beyond chasing price action, a key fundamental metric for evaluating DeFi protocols is sustainable revenue, which indicates real user demand. This analysis highlights high-revenue projects across key categories, using protocol fee data (net of supplier payouts). **DEX** * **Uniswap (UNI)**: Leads with $7.18M in 30-day revenue. Protocol fees from v2 and select v3 pools are used for UNI token burns. * **Solana DEXs**: Jupiter (JUP, $4.69M 30-day revenue) uses 50% of revenue for JUP buybacks. Meteora (MET, $1.67M) and Raydium (RAY, $1.13M) also allocate portions of fees to token buybacks. * **PancakeSwap (CAKE)**: Earned $5.16M in 30 days, with part of its fees used for CAKE burns, maintaining a net deflationary supply. * **Aerodrome (AERO)**: On Base, it generated $4.11M in 30 days. Revenue is directly distributed to veAERO holders rather than used for buybacks. **Lending** * **World Liberty Financial (WLFI)**: Top earner with $10.47M in 30-day revenue. A proposal passed to use 100% of fees from its Protocol-Owned Liquidity (POL) for WLFI buybacks, but holder net income remains zero. * **Aave (AAVE)**: Generated $4.12M in 30 days. Its buyback program was paused in April 2026 following the rsETH bridge attack. **ETH Staking** * **ether.fi (ETHFI)**: Earned $3.03M in 30 days. Revenue from eETH withdrawals is used for ETHFI buybacks, which are then distributed to sETHFI stakers. * **Lido (LDO)**: Generated $2.31M. Its new NEST mechanism automatically uses 50% of annual revenue exceeding $40M for LDO buybacks. In summary, several DeFi protocols are generating significant revenue, with many employing token buyback or direct distribution mechanisms. This revenue provides a fundamental basis for evaluation amid market volatility.

marsbit3m ago

DeFi Sector Bounces Back Strongest: Which High-Revenue Projects Offer Entry Opportunities?

marsbit3m ago

Hyperliquid is also getting a Layer2, what is Elysium?

Hyperliquid, a decentralized exchange, is set to launch its own Layer 2 solution called Elysium, developed by its largest liquid staking protocol, Kinetiq. This move aims to address key limitations in Hyperliquid's current ecosystem, particularly the performance and user experience issues on its existing HyperEVM. The article explains that HyperEVM has struggled with network congestion, high gas fees (sometimes exceeding $10-$20 per simple swap), and a fragmented infrastructure for launching and trading new tokens, especially memecoins. While there is significant speculative interest, the current setup lacks the efficient trading infrastructure to sustain it. Elysium is designed as a high-performance L2 that will use HYPE as its gas token. Its goals are to provide drastically faster block times and higher throughput compared to HyperEVM, create a seamless pipeline for token launches (from initial creation on Elysium to eventual listing as spot and perpetual markets on Hyperliquid's main chain, HyperCore), and offer developers richer access to HyperCore's order book data for better hedging and market-making. The L2 is positioned not as a competitor to HyperCore but as a "value-accrual" layer that aims to drive more activity and volume back to the main chain. Potential use cases extend beyond memecoins to include complex applications like PaperTrade (a novel perpetual DEX) and other DeFi protocols requiring fast settlement and real-time data. Elysium's sequencer revenue is planned to be shared with applications, the Kinetiq treasury, and used to buy back and burn the KNTQ token.

marsbit4m ago

Hyperliquid is also getting a Layer2, what is Elysium?

marsbit4m ago

Towns Across the U.S. Resist Data Center Construction, Bitcoin Miners Become the Biggest Beneficiaries

Across numerous U.S. states, governors and local governments are slowing or halting approvals for new data center construction due to intense public opposition and political pressure. This creates a significant hurdle for AI and other data center projects, which require years of grid connection approval. According to Morgan Stanley, this grid access bottleneck unexpectedly benefits Bitcoin mining companies like Cipher, Hut 8, and Riot Platforms. These miners already possess pre-approved, valuable power contracts and grid connections. With new supply constrained, their existing infrastructure becomes more valuable. This dynamic accelerates a pre-existing trend: miners are pivoting to supply power and data center capacity to the AI industry, signing large contracts like the $9.1 billion deal between Riot Platforms and Anthropic. Companies with AI/high-performance computing contracts command higher valuations than pure-play Bitcoin miners. However, the shift faces uncertainty. Regulatory reviews, like one in Texas targeting both data centers and crypto mines, could also impact miners' permits. Furthermore, a recent sharp rally in Bitcoin's price, pushing toward $80,000, could make mining profitable again, potentially altering the economic calculus of diverting power to AI. The core conflict pits growing AI demand against community resistance to large, power-intensive facilities.

marsbit6m ago

Towns Across the U.S. Resist Data Center Construction, Bitcoin Miners Become the Biggest Beneficiaries

marsbit6m ago

US Stock Market Trend (Aug 25): Besant's Trillion-Dollar Bond Buy Pushes Down Long-Term Rates, Nasdaq Falls, Dow Rises Amid Extreme Divergence

August 26, US Stock Market: Nasdaq Drops as S&P 500 Sees Mixed Close Amid Treasury Moves, Geopolitical Concerns. The US stock market was sharply divided on Monday. The Dow Jones rose 0.26% to 53,417.16, extending gains, while the S&P 500 slipped 0.28% to 7,652.82, and the Nasdaq Composite fell 0.76% to 25,869.14. The Philadelphia Semiconductor Index tumbled nearly 4%. Treasury Secretary Bessenter took center stage, announcing dual actions: deploying nearly a trillion dollars from the Treasury General Account (TGA) to repurchase long-term bonds to pressure yields, while simultaneously imposing new sanctions on Iran. Long-term Treasury yields edged lower, with the 10-year yield falling about 3 basis points to 4.70%. However, markets focused on the geopolitical risks from the sanctions, prompting a risk-off shift into gold and Bitcoin, weighing heavily on tech and AI-hardware stocks. Nvidia informed major cloud clients of plans to raise AI server prices by over 15% starting next year. The company also reportedly plans to invest in AI search firm Perplexity. Despite this pricing power signal, its stock and the broader chip sector were dragged down by macro and geopolitical worries. All eyes are on Nvidia's earnings report scheduled for release after Wednesday's market close, with focus on Blackwell shipment details and data center revenue guidance. Samsung Electronics disappointed markets with a shareholder return plan seen as insufficient, raising concerns about capital expenditure prospects in the memory chip sector and contributing to the semiconductor selloff. In trade policy, former President Trump announced plans to raise US tariffs on Canadian autos, auto parts, and steel to 50% starting next year, significantly escalating trade tensions. Amid the risk-off sentiment, spot gold rose 1.05% to $4,651.24 per ounce, hitting a near three-month high. Bitcoin climbed 1.59% to $78,966, approaching the $80,000 level for the first time since mid-May. Market attention turns to the US Conference Board's August Consumer Confidence Index, for clues on the resilience of consumer spending, and to continued positioning ahead of Nvidia's pivotal earnings report.

marsbit9m ago

US Stock Market Trend (Aug 25): Besant's Trillion-Dollar Bond Buy Pushes Down Long-Term Rates, Nasdaq Falls, Dow Rises Amid Extreme Divergence

marsbit9m ago

Trading

Spot
活动图片